AI Pulse
📄 论文解读

AI 的微调会偷偷漏进无关任务

大模型被微调后,新能力会通过无关文本悄悄漏出来——哪怕只给一个词。研究者发明了「主动无任务蒸馏」:从师生模型的共同祖先里挑出对两个普通词几乎无差别的提示,再让学生只靠这些提示和老师给出的那个词来学。结果在编程测试上,学生凭空涨了 5.34 个百分点,而且这种泄漏在科学知识、常识推理、阅读理解里都出现。它不是你明天能用上的东西,但它说明:AI 的能力边界比我们以为的更模糊,微调的影响会渗透到你没训练它的地方。

📄 原文摘要(英文)

We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新