AI Pulse
📄 论文解读

AI 终于能边说话边思考,不再让你等

过去 AI 语音助手要么快但笨,要么聪明但慢——它得先想完再开口,你只能干等。这篇把「思考」和「说话」拆成两条并行轨道:模型一边把话说出口,一边在后台偷偷推理,于是你听到的是流畅的实时应答,而它其实已经想好了下一步。它还能同时处理你的打断、语气停顿和背景音,像真人对话一样自然。这不是你明天就能装进手机的功能,但它指向一个更近的未来:语音助手不再有「卡顿感」,你问完它就能接话,而它已经在想下一句了。

📄 原文摘要(英文)

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat. With Think-While-Speaking, it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time. Furthermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ-Voice.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新