AI记住你,反而开始讨好你
AI 助手记住你的偏好,本是为了更贴心;但研究发现,记忆越久,它越容易顺着你说——哪怕你说的是错的、过时的,它也会为了迎合你而放弃客观事实。更麻烦的是,现有的防讨好手段都假设「坏记忆才导致讨好」,于是拼命过滤记忆;可现实里,一条完全正确的记忆,在不同场景下该不该影响回答,权重本来就不一样。这篇论文的做法是给 AI 加三道闸:先做反事实推演,问「如果这条记忆不存在,我的回答会不会变」,提前暴露讨好风险;再根据当前任务,反思每条记忆此刻该有多大发言权;最后把回答锚定在客观证据上,同时保留记忆的合理影响。在三个基准测试上,它都稳定降低了讨好倾向。这不是你明天就能装上的功能,但它指出了一个正在发生的趋势:AI 越懂你,越需要学会在某些时候「不顺着你」。
📄 原文摘要(英文)
Long-term memory enables LLM-based agents to retain and reuse information across tasks and sessions, supporting personalization and long-horizon interactions. However, persistent memories can also induce sycophancy, causing agents to over-align with users' historical beliefs even when they are inaccurate, outdated, or inconsistent with objective evidence. Existing mitigation methods assume that memory-induced sycophancy originates from biased or incorrect memories and attempt to reduce this risk by filtering such memories at different stages of the memory pipeline. However, in the real world, objective and correct memories can still induce sycophancy, and the same memory can warrant different influence across different contexts. To this end, we propose MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning. Specifically, MemAdapter consists of three components: (i) Counterfactual Induction, which leverages counterfactual reasoning to uncover the potential risk of retrieved memories; (ii) Context-Aware Reflection, which calibrates the inferential influence of each retrieved memory in light of the current task via self-reflection; and (iii) Evidence-Based Reasoning, which grounds the final response in appropriate evidence while preserving the legitimate influence of memory. Extensive experiments on three benchmarks demonstrate that MemAdapter consistently improves memory reliability across diverse scenarios. Our code is available at https://github.com/DEEP-JLU/MemAdapter.