AI Pulse
📄 论文解读

让两个AI互相抄作业,比各自刷题进步更快

训练AI数学推理有个尴尬时刻:模型自己生成解题过程,如果全错,就没有任何信号告诉它往哪改。这篇论文发现,两个不同的模型往往擅长不同的题——A不会的B会,B不会的A会。于是研究者让它们互相交换解题过程:A的题做不出来,就把B做对的答案拿过来当学习材料,连B算出的“这步有多重要”也一起带过来。为了防止抄错家作业(模型风格差异导致误判),系统给每条交换的轨迹按相似度加权、按词级裁剪,控制风险。在三个模型组合、五个数学基准上,两个模型都比各自单独训练涨分,平均涨2.1分,最多涨4.5分,而且存下来的交换轨迹不用同时训练也能保留大部分收益。这不是你明天能用的东西,但它指向一个趋势:AI训练不再只是单个模型埋头刷题,而是让不同模型互相当老师,成本更低、进步更快。

📄 原文摘要(英文)

Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新