大模型泛化被误解了:稳定才是真本事
我们一直用「准确率」衡量大模型会不会举一反三,但这篇论文说:准确率高不等于泛化好,稳定才算。同一个问题换个说法,模型答案可能就变了——这才是泛化的真相。研究者提出一套新框架,不只看答对多少,而是看同一输入在不同表达、不同任务下,模型的输出、内部激活、置信度有多大波动。结果发现:没有哪个模型在所有变化下都稳定,不同维度暴露的毛病互不相干,换个数据集甚至能让模型排名反转。也就是说,你平时看到的榜单排名,可能只是某个特定问法下的偶然结果。它不是你明天能用上的东西,但提醒你:别把跑分当人品,真正靠谱的模型,是换个问法依然不慌的那种。
📄 原文摘要(英文)
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.