AI推理提速10倍:新方法打破猜测解码天花板
大模型推理慢,一个常用加速技巧是「猜测解码」:让小模型先快速猜几个词,大模型再批量验证。但猜得越多,猜对的概率越低,而且猜的过程本身也耗时间,所以提速有天花板。JetSpec 把猜词过程改成「一次前向传播生成整棵候选树」,同时让每个分支的猜测都基于前一个分支的结果(因果条件化),这样猜出来的词更一致、更可能被大模型接受。在数学题上达到9.64倍加速,日常对话也有4.58倍。这不是你明天能直接用的工具,但它指向一个趋势:AI推理的瓶颈正在被系统性地拆解,未来更快的模型服务会来自这类底层优化。
📄 原文摘要(英文)
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and drafting overhead stays low. This ceiling has been difficult to break because prior head-based SD methods face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates that are effective for tree speculative decoding with higher acceptance length, but their drafting cost grows with tree depth. Bidirectional block-diffusion drafters generate all positions in one pass, but their branch-agnostic marginals can form individually plausible yet mutually inconsistent trees, wasting budget and reducing acceptance. We propose JetSpec, a head-based SD framework that combines one-forward drafting efficiency with branch-wise causal conditioning. JetSpec trains a causal parallel draft head over fused hidden states from the frozen target model, producing candidate trees whose scores align with the target model's autoregressive factorization. This enables JetSpec to convert larger draft budgets into longer accepted prefixes and higher end-to-end speedup. Across math, coding, and chat benchmarks on dense and MoE Qwen3 models, JetSpec consistently outperforms bidirectional-head and tree-based SD baselines. On H100 GPUs, JetSpec achieves up to 9.64x speedup on MATH-500 and 4.58x on open-ended conversational workloads, with further latency gains demonstrated through vLLM integration under realistic serving loads. Our code and models are available at https://github.com/hao-ai-lab/JetSpec.