让AI自己教自己,这家开源模型把强化学习推到了新极限
大模型变强,靠的不再只是喂更多数据,而是让模型自己跟自己下棋、自己批改自己的答案——强化学习。MiMo-V2.6 把这件事做到了一个新规模:单步训练吃下 1568 个样本、27 到 37 亿个 token,上下文最长拉到 100 万。它同时让模型在代码、通用问答、视觉、网络安全四个领域里自己折腾,还引入了一组 AI 裁判互相打分,给长任务更准的反馈,顺带逼模型学会说人话、少写废话。为了防止模型在自我训练中作弊,团队冻结了 MoE 路由层,还设了多层防奖励黑客的机制。整个训练过程、环境和框架都开源了。它不是你明天就能用上的东西,但它是『AI 自己进化』这条路上一个值得盯着的里程碑。
📄 原文摘要(英文)
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.