AI 开始参与造自己的下一代,人类还剩什么
这篇论文的主角不是某个新模型,而是一个正在发生的分工转移:AI 不再只是被人类使用的工具,它开始参与设计自己的继任者。研究者训练了一个叫 Atria Dawn 的模型,专门用于科研和工程,并记录了 56 个人用它做 769 个任务的全过程。结果发现,AI 频繁提出方法和修改方案,但人类保留最终决定权——参与者自己评估,约三分之一的任务没有 AI 根本做不成。也就是说,人类的重心从「怎么做」挪到了「值不值得做、证据该怎么用」。这不是你明天能上手的东西,但它指向一个更值得盯住的趋势:AI 的进步不再只是能力竞赛,更是人类监督能力的竞赛。
📄 原文摘要(英文)
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.