AI Pulse
📄 论文解读

给AI装上左右脑:一边记事实,一边懂情绪

现在的语音AI能听懂你说什么,却记不住你是谁——它不知道你上周提过妈妈住院,也听不出你声音里的疲惫。VoiceMem给语音模型装了个双脑记忆:左脑管事实,像搜索引擎一样快速调出你之前说过的关键信息;右脑管情绪和人格,记住你说话时的语气、你的性格偏好。实测里,它调取记忆只要134毫秒,比人眨眼的十分之一还短,完全不影响对话节奏;记事实的准确率比现有最好的记忆系统高出一大截,懂情绪和人格的能力也在三个标准测试里拿了第一。它不是你明天就能用上的东西,但这是语音助手从“应答机器”变成“记得你的人”的关键一步。

📄 原文摘要(英文)

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新