AI Pulse
📄 论文解读

AI 学不会的题,靠「事后复盘」学会了

强化学习有个死穴:如果一组尝试全失败、得分全是 0,AI 就学不到任何东西,直接卡死。这篇论文给了一个反直觉的解法——让 AI 在行动之前,先「预演」一遍自己会怎么做,然后拿真实行动后的结果去纠正这个预演。相当于考试前先写一遍答案草稿,考完对答案后回头改草稿,下次考试草稿就更准。在 2B 规模的模型上,原本 0% 成功率的任务,加上这套「自我复盘蒸馏」后达到 60.6%,而且不增加任何推理时的计算成本。它不是你明天能直接用的功能,但它指向一个趋势:AI 的学习不再只靠「对错」,也开始靠「复盘」——这可能是下一代模型突破瓶颈的方向。

📄 原文摘要(英文)

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp. Its advantage is especially pronounced when reward contrast is scarce: when 37--98% of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where 98% of groups are all-failure, the RLVR training ends up at 0.0% success, while adding SRD reaches 60.6% under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新