AI离自主科研还差得远:撤掉指导,能力暴跌
一个由40多位专家、耗时3.1万小时建成的基准测试,给当前最强AI做了一次“放手”实验:当人类把研究方法也撤掉,让AI自己决定怎么做科研时,平均得分从50.91掉到26.62,几乎腰斩。这印证了一个反直觉的事实——今天AI的“聪明”高度依赖人类喂给它的步骤,一旦没人指路,它连怎么开始都犯难。这不是说AI没用,而是说它离“超级智能”的想象还很远:它擅长在既定框架里做到极致,却不擅长自己开辟新路。对普通人而言,这其实是个好消息:短期内,AI更像一个需要你不断给指令的强力工具,而不是能独立搞发明的“外星大脑”。
📄 原文摘要(英文)
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.