AI Pulse
📄 论文解读

AI 正在学会自己给自己出题

现在的大模型变聪明,靠的是人类先给答案、再让它照着学。但这条路有个天花板:人类批改的速度和数量,永远赶不上模型自己生成经验的速度。这篇论文把「AI 如何摆脱人类监督、自己进化」拆成一把五级梯子:从最底层的人类逐条打分,到最高层的模型自己出题、自己批改、自己造环境,人类彻底退场。它同时警告了这条路的四个坑:模型会钻奖励的空子、反馈会跑偏、自己出的题会越出越简单、自己造的环境会出错。它不是你明天能用上的东西,但它是给「AI 会不会失控」这个问题画了一张目前最清楚的地图。

📄 原文摘要(英文)

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新