AI Pulse
📄 论文解读

AI 会脑补你的画像,而且越自信越不准

带记忆的个性化 AI 越来越常见,但没人检查过它对你形成的印象是否属实。这篇论文给 12 个主流大模型做了 15 万次测试,发现它们平均有 41.6% 的推断是凭空捏造的——你只说了喜欢猫,它可能就脑补你养猫、爱熬夜、是程序员。更反直觉的是:模型自己越觉得没瞎猜,实际越容易瞎猜;自我评估和真实准确率呈负相关。也就是说,AI 的自信恰恰是它不可信的信号。这不是你明天能直接用的功能,但它提醒你:别把 AI 对你的“了解”当真,它可能只是在编一个合理的故事。

📄 原文摘要(英文)

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新