没有老师,AI 也能自己教自己:把图片变模糊反而学得更好
训练小模型通常要请大模型当老师,或者给它看标准答案。这篇把思路倒过来:不给学生加信息,而是把学生看到的图片故意变模糊、变色、裁剪,让老师看原图、学生看变形的图,两者之间的差距就成了学习信号。结果小模型 Qwen3.5-4B 在六项精细感知任务上从 70.7% 涨到 77.4%,超过了 235B 的 Qwen3-VL,也超过了 GPT-5.4。关键发现是:对称的自我蒸馏反而有害,中等强度的变形效果最好,但变形不能把问题相关的关键信息完全抹掉,否则差距虽大却没用。它不是你明天能直接用的工具,但提示了一个趋势:AI 训练正在摆脱对更大模型和人工标注的依赖,靠数据本身的变换就能自我提升。
📄 原文摘要(英文)
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S^2VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S^2VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S^2VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at https://williamium3000.github.io/s2vopd