去掉视觉编码器,大模型反而能追上?
现在多模态大模型几乎都靠一个预训练好的视觉编码器来「看懂」图片,相当于先请个专家把图像翻译成模型能懂的语言。这篇论文想验证:如果把这个专家撤掉,让模型直接从原始像素学,会不会更简单、甚至更强?他们在不同算力规模下对比了两种架构,发现一个反直觉的结果:小规模时,没有编码器的模型明显更差;但按缩放规律推算,当训练算力达到约 10^22 FLOPs(在主流预训练预算内),它就能追平甚至反超。更关键的是,模型自己会「长出」视觉能力:随着算力增加,视觉 token 之间的交互越来越重要,处理视觉信息的计算会主动移到更早的层,甚至出现专门处理图像的「专家路由」。换句话说,人类精心设计的视觉先验,在足够大的算力和数据面前,优势会被抹平。这不是你明天能用的技术,但它指向一个趋势:多模态模型可能走向更统一的架构,少一个组件、少一层依赖。
📄 原文摘要(英文)
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around 10^{22} FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.