AI模型瘦身:只留关键层,长文本处理快10倍
大模型处理长文本时,全注意力机制计算量巨大。以往的做法是随机或凭经验保留部分注意力层,但效果不稳定。这篇论文把问题转化为数学优化:先给每个注意力层加一个“线性注意力”分支,然后冻结模型权重,只训练一个开关门控,让模型自己决定哪些层该保留全注意力、哪些可以简化。最终在保持长文本召回率的同时,大幅降低计算成本。它不是你明天就能用的工具,但为模型部署提供了更聪明的剪枝思路。
📄 原文摘要(英文)
Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid conversion critically depends on which layers preserve full attention. Existing hybrid layer selection methods typically rely on heuristic strategies such as fixed placement patterns or layerwise scoring, implicitly treating layer importance as isolated and overlooking the interdependent layer effect under a global hybrid configuration. In this work, we formulate hybrid layer selection as a budget-constrained subset optimization problem. We further propose FlashMorph (Fast LAyer Selection for Hybrid MORPHing), an effective, efficient and scalable layer selection method for Transformer-to-hybrid conversion. FlashMorph first constructs a morphable model by equipping each full-attention layer with a converted linear-attention branch. It then freezes all model weights and jointly optimizes layerwise gates on synthetic long-context retrieval data, with a linearization regularization that encourages the model to rely on linear attention for efficiency. The learned gates are discretized under a preset full-attention budget to instantiate the hybrid architecture, followed by standard logits distillation and long-context finetuning. Extensive experiments show that FlashMorph discovers more effective hybrid configurations, preserves strong long-context recall and general benchmark performance while substantially reducing layer selection cost compared with existing layer selection methods, demonstrating its effectiveness, efficiency, and scalability.