大模型有周期性失忆点
给大模型压缩记忆缓存,本意是省内存、加快长文本推理。研究者发现这招会埋下规律性盲区:同一段信息,放在压缩窗口的某个位置一查就中,挪到另一个位置就死活想不起来,准确率能差 40 个百分点,而平均分完全看不出来。他们从零训练了一族模型,确认这不是偶然,还定位到是不同注意力组件对不同位置各管一段,才造成这种周期性偏科。它不是你明天能用上的东西,但提醒你:以后看长文本 AI 的评测分数,得留个心眼——高分可能只是没考到它的盲区。
📄 原文摘要(英文)
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.