AI Pulse
📄 论文解读

让AI变强,不一定用更强的AI来教

训练AI时,通常用大模型生成“错误答案”来教小模型,认为错误越像自己犯的越有用。但研究发现,用更小的模型生成错误答案,反而能训练出更强的学生模型,而且计算成本更低。错误答案的价值不在于“多像”,而在于保留任务结构、与参考策略保持距离。这推翻了“越大越好”的直觉,也意味着未来训练AI可以更便宜、更高效。

📄 原文摘要(英文)

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新