大模型嘴上说得很自信,但它的“怀疑”藏在哪一层?
大模型经常用流畅的话和笃定的语气说出错答案,你没法从字面判断它是否在胡编。这篇论文不靠重复生成或额外训练,而是直接钻进模型的内部状态,找出它“怀疑”和“确信”分别对应的神经信号方向,做成一把尺子,逐字量出模型在推理过程中哪一步开始心虚。它给出的置信度比现有方法更准,而且不是靠“答案长短”这种表面特征蒙混过关。它不是你明天能用上的工具,但它把“AI 什么时候在硬撑”从玄学变成了可观测的物理量。
📄 原文摘要(英文)
Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: https://github.com/s2labres/U-Space.