AI推理省内存,不再需要手动调参数
大模型推理时,KV缓存(存储已处理信息的临时记忆)会吃掉大量显存。现有压缩方法虽然能省内存,但需要你提前告诉它“留多少缓存”——这个数字得针对具体输入手动调,换一个领域或任务就失效,否则性能暴跌。这篇论文直接砍掉这个调参步骤:让模型自己根据输入动态分配缓存预算,在13个不同数据集上都能保持和完整缓存一样的准确率,同时省下大量内存。它不是你明天就能用的工具,但指明了方向:未来的AI压缩应该自适应,而不是依赖人工预设。
📄 原文摘要(英文)
To reduce memory consumption during LLM inference, a handful of methods have been proposed for KV cache pruning. While these techniques can accomplish lossless memory reduction on many datasets, they often hinge on an under-emphasized condition: an input/domain-specific threshold for KV cache budget needs to be pre-determined to achieve the optimal performance. However, such input-sensitive design may be considerably limited in real-world scenarios, as open-domain inputs span diverse domains, lengths and difficulty levels, without clear boundaries for threshold selection. As a result, the dependence of such input-sensitive threshold can be a fundamental limitation that causes large degradation on arbitrary inputs. In this work, we propose a new objective that lifts the threshold constraints for robust KV compression, advocating for "threshold-free" methods that adaptively adjust budget allocation while preserving full-cache performance. We then propose a novel method, ReFreeKV, serving as the first instantiation of this objective. Extensive experiments across 13 datasets with diverse context lengths, task types, and model sizes demonstrate its efficacy and efficiency. Our code is publicly released at https://github.com/Patrick-Ni/ReFreeKV.