AI记忆不用提炼,直接挑原文就够了
给AI配记忆,主流做法是把对话先提炼成摘要再存起来,费钱又费时。这篇预注册研究在LoCoMo和LongMemEval上发现:直接用一次调用挑出最相关的原始对话片段,效果不输给提炼式记忆,而且写入成本便宜3061倍。预算紧时,重排序的增益会缩水到1.5分,提炼系统反而更准——这解释了为什么之前的研究结论打架。它不是你明天能用上的,但提示了一个方向:AI的记忆可能不需要那么聪明,挑对原文就够了。
📄 原文摘要(英文)
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.