把约束从答案挪到问题,RL 训练稳了
训练大模型时,我们一直靠「别让回答跑太远」来稳住优化,但这会牺牲探索新答案的空间。阿里这篇换了个思路:不去管回答,而是管「问题」——训练中模型会越来越偏爱你喂给它的那些题,他们加了一项约束,让模型对问题的偏好别偏离原始分布太远,同时完全不动回答的探索自由度。在 6 个数学推理基准上,它替换掉原来的约束,准确率更高,高温解码和长训练下也更稳。它不是你明天能用上的东西,但提示了一个方向:很多 AI 训练里的两难,可能只是你把约束放错了位置。
📄 原文摘要(英文)
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in a double bind: keeping Policy-KL constrains response behavior and consumes the action-side exploration budget, while dropping it leaves the optimization without an explicit drift control. We argue for an alternative that breaks the dilemma by moving regularization to the input side. As training progresses, the distribution over training queries induced by the current policy drifts unchecked from its pre-RL reference distribution. Concretely, Environment-Regularized Policy Optimization (ERPO) introduces a Query-KL (QKL) term that bounds this query distribution shift, together with a dataset-static reference-derived per-query weight that biases each per-query update toward queries typical under the reference. The QKL gradient flows strictly through the query likelihood; the response score function used by policy-gradient estimators does not appear in the QKL term, so QKL exerts no direct gradient pressure on the response distribution---exploration is preserved. ERPO plugs into GRPO/PPO/REINFORCE-style pipelines without additional forward passes. On six mathematical reasoning benchmarks, ERPO replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.Our source code are available at https://github.com/alibaba/ERPO