让AI自己改自己,改到第几轮才见真章
我们总以为AI的聪明是训练出来的,这篇告诉你,聪明也可以是“自己改”出来的。研究者给AI布置了机器学习、算法编程这类能自动判对错的题,让它反复试错、反复修改,练出一套“越改越好”的本事。练完的AI,在数学、编程、甚至深度搜索任务上都变强了,而且给它更多轮次去改,成绩还会继续涨。它不是你明天就能拿来用的工具,但它指了个方向:AI的自我进化,可能不需要更聪明的模型,只需要更聪明的“练习方式”。
📄 原文摘要(英文)
We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.