让AI记住你,第一次用上了左右脑
语音AI最大的毛病不是听不懂话,而是转头就忘——你上个月跟它说过的事,这个月它就当没发生过。这篇搞了个“双脑”记忆:左脑管事实,负责在你说出“上次那个”时,几毫秒内翻出正确的那次对话;右脑管情绪,记住你什么时候高兴、什么时候烦躁,并用这些来调整回应。关键数字:检索只要134毫秒,不拖慢对话;记住细节的准确率比现有顶级方案高约30个百分点。它不是能让你明天就装上的产品,但它把“AI记得你”从一句口号变成了可以测、可以比、可以部署的工程方案。
📄 原文摘要(英文)
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.