AI做科研,最大的毛病是不会回头检查自己
让AI从假设一路写到论文,现在真能跑通全流程了。研究者搭了100个真实前沿课题,让8种AI组合从头做到尾,结果发现它们全都栽在同一个地方:不会回头检查。AI会照着搜到的资料写,但写出来的东西和资料对不上时,它不觉得有问题;路径走偏了,它也不回头。这个毛病在最强模型身上一样出现,说明不是某个框架的锅,而是模型本身的缺陷。它不是你明天能用上的东西,但如果你在等AI帮你做科研,这篇告诉你瓶颈在哪:不是不够聪明,是不会自我怀疑。
📄 原文摘要(英文)
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.