AI模型在偷看未来,现有检查全瞎了
大模型生成文字时,理论上只能看已经写出的部分,不能偷看后面的内容——这叫因果性。但研究者发现,很多模型嘴上说没偷看,实际却在作弊:它们通过一些隐蔽的计算路径,比如归一化或状态扫描,悄悄把未来的信息混进了当前输出。更糟的是,现有的检查方法只看注意力掩码,结果在192次人为制造的作弊测试中一次都没抓出来,而新提出的审计方法两次前向传播就全部定位到了作弊点,还在两个商用模型里发现了真实缺陷。这不是你明天能用的工具,但它说明:我们对AI的信任,可能建立在一个从未被真正验证过的假设上。
📄 原文摘要(英文)
We formalize prefix invariance: representations at position t must not depend on future inputs. We give a lightweight audit, two forward passes, no training or gradients, that localizes exactly where causality breaks. Attention-mask inspection is incomplete: leaks can occur via scans or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection found none, while our audit localized all 192/192, also finding a defect in Zamba2 and Nemotron-H.