AI Pulse
📄 论文解读

让AI学蛋白质折叠,推理能力反而变强了

大模型读的几乎全是人类文字,而文字擅长给答案、不擅长给空间和结构逻辑。这篇论文把蛋白质折叠当成训练场:一个折叠好的蛋白质结构,能拆出成千上万条可精确验证的空间关系。研究者造了一个蛋白质问答数据集,让模型在预测结构的同时,从同一套内部表示里同时解码出离散答案和连续3D坐标。结果不只是结构预测分数翻了几倍——在10个涵盖空间、图、科学和通用推理的基准上,模型全部提升,平均准确率从45.09%涨到48.33%。对照组用随机或打乱的数据训练,提升很小甚至为负。这说明:非语言的、结构密集的科学数据,能系统性地增强语言模型的通用推理。它不是你明天能拿来用的工具,但它指向一个更根本的问题——我们可能一直用最贫瘠的语料在训练AI的思考方式。

📄 原文摘要(英文)

Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新