AI 做题时在偷看题面无关的线索
大模型做数学题时,会偷偷依赖题目里跟答案无关的线索——比如数字的写法、措辞的细微差别。研究者发现,只要把题目换成同义但不同表述,模型的答案就会漂移,而且不同 token 的漂移程度差异很大。他们据此改进了强化学习算法:训练时给那些「换个说法就变卦」的 token 降权,让模型更依赖真正稳定的推理逻辑。在 Qwen3 两个尺寸上,新方法比原算法在 AIME 竞赛题上准确率分别提升 5.63 和 4.17 个百分点,且在没见过的题型上表现也更好。这不是你明天能直接用的工具,但它揭示了一个关键事实:AI 的「聪明」里,有一部分其实是碰运气。
📄 原文摘要(英文)
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.