AI Pulse
📄 论文解读

AI陪聊三个月,到底懂你多少?

AI 陪聊产品都在吹「懂你」,但没人敢拿真实聊天记录来测——隐私太重。这篇论文反着来:公开了 10 段真实用户和 AI 的长期对话,共 27,218 条消息、最长 120 天,并附上逐条标注的「记忆点」和推理过程。结果很打脸:绝大多数时候,AI 根本不需要回忆过去——95.9% 的问题靠最近几条消息就能答,真正需要翻旧账的只有 2.2%;而且现有模型分不清「什么时候该用记忆」,一旦把同一句话标成「记忆」,模型使用它的概率立刻涨 10 到 14 个百分点。更讽刺的是,三个主流 AI 系统重建用户画像的准确率一样,但成本差 31 倍。这不是你明天能用的功能,但它给「AI 懂你」这个营销词划了条诚实的底线:目前 AI 的「懂」,更多是即时反应,不是长期理解。

📄 原文摘要(英文)

A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 120 days, released as the conversation and four files derived from it, a profile, a persona, a chat ground truth and a question set, each citing the messages it rests on. Every chat label carries the reasoning trace that produced it, checked stage by stage against the conversation. Three findings follow. First, the past is rarely needed and far away. Pooled measures mislead: a recency window finds the required message for 95.9\% of probes and 2.2\% of those that need memory, and at the natural rate 96\% of the gain from supplying recorded evidence comes from messages that need none. Second, no detector we tried can tell when memory is needed on real messages, authored questions over the same histories leak the cue, and labeling the same messages as memories raises their use by ten to fourteen points. Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新