AI Pulse
📄 论文解读

剪掉一半专家,推理反而快一倍

大模型推理时,真正卡脖子的不是算力,而是把专家权重搬进显存的数据流量。以往剪枝为了省流量,连负责理解输入的 prefill 阶段也一起剪,结果模型变笨、速度却没快多少。SlimWise 反着来:理解阶段用完整模型,生成阶段才用剪掉一半专家的精简版,而且两阶段之间共享 KV 缓存、无需转换。在 Qwen 等模型上,剪掉 50% 专家后,生成吞吐最高提升 1.81 倍,准确率损失很小。它还发现一个坑:基准测试的准确率看着没掉,但生成长度悄悄变了——这提醒我们,只看分数可能被误导。这不是你明天能直接用的工具,但它指向一个趋势:未来的推理优化不是一刀切,而是按阶段精细调配资源。

📄 原文摘要(英文)

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新