AI Pulse
📄 论文解读

小模型当老师,能教出比自己强的大模型

强化学习能让大模型变聪明,但这份聪明能不能传给别的模型、传多少,一直没人说清。这篇论文发现一个反直觉的规律:在「弱教强」的师生组合里,学生最终能达到的巅峰水平,普遍超过老师自己。也就是说,一个算力有限的小模型,经过强化学习练成专家后,可以把能力蒸馏给大得多的模型,而且学生能青出于蓝。研究者还拟合出幂律,预测巅峰成绩和师生规模的关系:老师变大带来的收益,到学生规模附近就到顶了;而两个老师如果分数一样,小的那个反而教得更好——所以老师的分数不能代表它的教学价值。这不是你明天能直接用的技术,但它给「用小模型训练大模型」这条省钱路线提供了理论依据。

📄 原文摘要(英文)

*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, G) rises approximately linearly in d=mathrm{KL(π_θVert π_{ref})}, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how G_{peak} and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新