AI Pulse
📄 论文解读

把图像压缩成文字一样的代码,还能不丢细节

AI 处理图像时,通常把图像变成连续的高维向量,这很占算力。这篇论文反其道而行,把图像压缩成离散的代码——就像文字一样——但以往这么做会丢失大量细节。ViQ 用两阶段训练:先让视觉编码器从语言模型那里学语义,再通过一种渐进式压缩和位置感知的量化机制,把图像变成紧凑的代码,同时保留低层细节。结果:在多个多模态任务上,ViQ 的性能不输传统连续向量方法,但训练速度提升了 20% 到 70%。它不是你明天能用上的,但它指向一个趋势:未来多模态模型可能不再需要高维视觉特征,而是像处理文字一样处理图像,更轻更快。

📄 原文摘要(英文)

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete signals in the same way as text inevitably introduces severe information loss. Existing work struggles to balance low-level details and high-level semantics in discrete representations: reconstruction-oriented representations often lack semantic information, whereas semantically stronger features typically suffer from severe loss of detail. We present ViQ, a Visual Quantized Representations framework, which is designed to balance semantics and details in discrete representations while supporting inputs at native resolutions, thereby enabling it to serve as a unified and general discrete representation for arbitrary visual inputs. Our approach structures quantization learning into two stages: text-aligned pre-training and feature discretization. With text-aligned pre-training, we enhance the visual encoder semantic-rich supervision from the pretrained language model and enable it to process native-resolution visual inputs. During discretization, we propose a proximal representation learning strategy to progressively compact the feature space, along with a position-aware head-wise quantization mechanism that enables flexible processing of arbitrary resolutions. Extensive experiments on multimodal tasks demonstrate that ViQ achieves competitive performance compared to state-of-the-art multimodal vision encoders with continuous and high-dimensional visual features, while maintaining high precision in low-level reconstruction. We also show that multimodal training with visual quantized representations largely improves efficiency, yielding up to 20\%-70\% acceleration with different base LLMs and training recipes.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新