AI 自己教自己,靠的是互相打分
现在的 AI 变聪明,主要靠人类给标准答案当“教练”。但问题来了:当 AI 的推理能力超过人类能可靠评判的水平,标准答案会越来越稀缺、越来越贵。这篇论文换了个思路:让一群 AI 互相打分、互相教,不靠任何人工标注。具体做法是让多个互不相干的模型(不共享参数)组成“班级”,每个模型用同伴的反馈当奖励来学习。关键发现是:班级越“杂”越好——模型类型、大小、训练样本的表述方式越多样,互相纠正错误的效果越强,还能防止所有模型退化成一模一样的“复读机”。在纯文本和图文混合任务上,这套方法比无监督基线平均提升 2.3% 到 8.6%,甚至追平或超过需要人工标注的监督方法。它不是你明天就能用上的工具,但它指向一个更省钱的未来:AI 的进步不必永远依赖人类当裁判。
📄 原文摘要(英文)
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self-generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi-agent training. We introduce Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse. Across text-only and multimodal domains, Co-RL consistently outperforms the base models and prior label-free approaches, while matching or surpassing supervised methods, without access to any ground-truth labels. Concretely, Co-RL yields average gains of 3.0-8.6% across seven text-only benchmarks for LLMs and 2.3-7.2% across four multimodal benchmarks for VLMs. Code is available at https://github.com/DrStranded/Co-RL.