拔掉大模型后门,不用重训,还不伤能力
大模型被植入后门后,现有防御手段往往像动手术:把后门切了,顺带把正常功能也切坏了。这篇提出一个不用重新训练的办法:先定位后门在模型内部对应的「方向」,再把这个方向上的权重掰正,同时专门保护跟「拒绝回答」相关的区域不动。在多个模型和攻击类型上,它把攻击成功率压到最低,代码注入攻击直接归零,而且对正常回答的影响比现有方法都小。它不是你明天能用上的工具,但「不重训、不伤能力」这个方向,是防御后门真正能落地的关键一步。
📄 原文摘要(英文)
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.