AI Pulse
📄 论文解读

AI 自己出题考自己,越考越强

现在的 AI 变强,大多靠人类喂数据、给反馈。这篇换了个思路:让 AI 自己当出题人、考生和考官,三个角色互相逼着进步——出题人专挑难题,考生硬着头皮答,考官不用自己打分,而是按「谁答得更好」的既定顺序来判。结果在能验证对错的领域平均涨 4.2 分,在没法验证的领域涨 8 分,而且练了十轮还在涨,对比之下,传统方法两轮就退步了。它不是你明天能用上的东西,但它指向一个方向:AI 可能不再需要人类手把手教,自己就能把自己练强。

📄 原文摘要(英文)

Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新