审计AI只需看5%的“话”就能发现威胁
给大模型做“安全审计”时,解释它每个内部状态太贵了。这篇论文发现:不用看全部,只挑5%最关键的“词位”去解释,就能保留几乎全部的威胁检出率——而且挑位置这件事,甚至不需要跑模型,只看对话结构就行。研究者用470万个解释,在提示注入和隐藏信息两类威胁上验证了这一点。它还发现,预训练好的“翻译器”能直接读出模型通过微调学会隐藏的词,不用重新训练。对做AI安全审计、红队测试的人来说,这意味着审计成本可以砍掉95%,而且更省事。
📄 原文摘要(英文)
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just 5% of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.