AI Pulse
📄 论文解读

大模型推理被拆成两半,各自用最合适的压缩方式

大模型回答你之前,其实有两段完全不同的苦力活:先把你整段问题一口气读完(叫 prefill),再一个字一个字往外蹦答案(叫 decode)。过去压缩模型时,这两段被迫用同一套方案,总有一边吃亏。这篇论文把两段拆开,分别定制压缩:读问题那段用低精度计算加速,蹦答案那段用更紧凑的权重减少内存搬运。在 Qwen 3 和 Gemma 3 上,只对 decode 去掉激活量化,就能在 decode 重的任务上提升准确率,还不增加推理成本;单独训练一套 prefill 专用权重,在 2-3 位压缩下,速度比纯权重推理快,准确率不降甚至更高。更狠的是,在 27B 模型上,把 prefill 权重放到 SSD 里流式读取,8K 长度提示的首字延迟比原来快 1.78 倍。它不是你明天能用上的东西,但说明一个趋势:大模型优化正在从「一套方案打天下」走向「按阶段分工」,就像 CPU 和 GPU 各干各的活。

📄 原文摘要(英文)

Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新