AI Pulse
📄 论文解读

AI 教 AI,教的是思路不是答案

大模型互相教,教出来的不是「这道题怎么答」,而是「遇到这类问题怎么想」。研究者把「学生模型自己出题、老师批改」的蒸馏方式拆开看:训练题的难度几乎不影响效果,甚至老师自己都答不出的题,也能让学生变强。真正决定学生能走多远的是「出身」——同一个家族出来的老师,能把推理能力带到别的语言、别的领域;不同家族的老师,学生就只学会了死记硬背。更麻烦的是,把多个专家老师混在一起教,学生的能力会像跷跷板一样此消彼长,你没法让每个老师只管自己擅长的部分。这不是你明天能用的技巧,但它解释了为什么「让 AI 教 AI」有时灵、有时不灵。

📄 原文摘要(英文)

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新