AI 终于要学「什么时候闭嘴」了
现在的语音 AI 只会抢话:你正跟朋友聊天,它突然插一句,或者该它回答时又装死。这篇论文给 AI 出了一套「多人饭局」考题——三个真人加一个 AI 同桌,看它能不能分清什么时候该答、什么时候该闭嘴、什么时候该停。结果最强的开源模型 MiniCPM-o 4.5 也只拿下三项中的领先,而其他模型要么话多但答错,要么该安静时还在说。更扎心的是:真人明确点名 AI 时它才反应,不点名就装没听见——但真人之间说话本来就不总点名。它不是你明天能用上的,但这是 AI 从「对讲机」进化成「饭搭子」必须过的一关。
📄 原文摘要(英文)
Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.