AI 自己给自己出题,越练越强
让 AI 变强,通常靠人类精心设计训练数据;这篇让 AI 自己当自己的老师。研究者用一个基础模型,在多种辅助环境下尝试解题,把成功的路径提炼成可复用的「操作手册」,再让模型照着手册在全新环境里重做一遍,把 2001 条成功经验扩写成 11094 条训练数据。练完以后,模型在终端操作类任务上的通过率从 57% 涨到 74%,在更难的测试集上从 1.5% 涨到 9.1%。它不是你明天能用上的东西,但它指向一个趋势:AI 的进步越来越不依赖人类喂数据,而是靠自己在复杂任务里反复试错、自我复盘。
📄 原文摘要(英文)
Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.