大模型训练只挑关键几步,200道题就够
训练大模型通常要喂海量数据,但一篇新论文发现:在扩散式语言模型里,真正决定答案走向的只有少数几个“关键承诺”——就像下棋时几步妙手定胜负,其余都是顺水推舟。研究者只挑出这些高影响步骤来训练,用200道题、每道4次尝试,就让模型在数学和代码任务上超过了传统全序列训练和强化学习的基线。它不是你明天能用上的东西,但提示了一个方向:AI训练可能不需要“多”,而需要“准”。
📄 原文摘要(英文)
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.