AI Pulse
📄 论文解读

大模型微调的秘密:方向比内容更重要

大模型训练有个怪现象:同样是用高质量数据教,“边做边学”的强化学习比“照着标准答案背”的监督微调学得更扎实、更会举一反三。这篇论文找到了一个关键差异:监督微调让模型的参数沿着一个固定方向走,而强化学习会不断调整方向。研究者试着把强化学习的方向“借”给监督微调——不让它自由发挥,只准它沿着强化学习指出的方向更新参数。结果,监督微调的泛化能力一下子追上了强化学习。这意味着,模型学得好不好,可能不取决于你喂了多少数据,而取决于参数更新的方向对不对。对普通人的启示是:AI 能力的提升,未必靠堆数据,可能靠更聪明的“走法”。

📄 原文摘要(英文)

The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新