AI Pulse
📄 论文解读

AI概念藏在角度里,不是长度里

我们一直以为大模型里「概念」既藏在向量方向也藏在向量长度里,但新研究用实验拆开两者后发现:概念几乎只由方向(角度)决定,长度(范数)不携带概念信息,却影响你操控模型时的稳定性。研究者用7个模型验证,并指出当前流行的「加一个向量」的操控方式其实同时改了角度和长度,导致效果不可控。它不是你明天能用上的,但解释了为什么有时调模型像调收音机——拧对了方向才有效,拧大了音量只会爆音。

📄 原文摘要(英文)

Linear activation steering has gained popularity as a simple and empirically effective way to control language model behavior. More recently, spherical steering paradigms have been proposed to address limitations of additive interventions, often motivated by the assumption that hidden-state norm does not carry concept-relevant information. In this work, we revisit this assumption through a controlled empirical study designed to disentangle the roles of angular and radial components. We show that steering methods differ mainly in how they couple two geometric effects: changing a token's angular alignment with a concept direction and changing its hidden-state norm. Across seven language models, we find that concepts are represented primarily in angular structure, supporting the motivation for spherical methods, but that norm remains important for the stability and downstream effects of steering. Our results explain why interventions with similar concept-level effects can behave differently, and suggest that activation steering should be parameterized by interpretable angular and radial components of the intervention, rather than by a single additive coefficient that entangles these two effects.

arXiv 原文

📬 订阅 AI Pulse

每天三次更新,不错过重要信号

▲ 回到顶部