把多个AI专家合成一个,关键不是选谁教,而是教多用力
把多个AI专家合成一个全能模型,直觉是让最擅长那道题的老师来教。但新研究发现:选对老师只解决了一半,另一半是控制每个老师的“嗓门”。研究者让数学、编程、指令跟随三个专家轮流教同一个学生,结果学生被“嗓门最大”的指令跟随专家带偏,数学专家的优势几乎没学到。他们提出的解法很简单:按每个领域反馈的离散程度重新调音量,数学反馈更集中就放大,指令反馈更分散就调小。在六个公开基准、三种模型规模下,这个调整稳定提升平均分,并找回大部分丢失的数学能力。这不是你明天能用的工具,但它揭示了一个反直觉的规律:AI融合的瓶颈往往不是“谁更懂”,而是“谁的声音更容易盖过别人”。
📄 原文摘要(英文)
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.