AI Pulse
📄 论文解读

AI 终于能边想边听了

现在的 AI 助手都是「一问一答」:你说话,它听完才想,想完才回。但真实世界不是这样——你正说着话它就该开始理解,它思考时你又有新消息进来。这篇论文把大模型改造成「异步」的:让模型能同时处理多个任务,边看视频边听指令、边打游戏边响应监控,而且不需要针对每个任务单独训练。它不是你明天就能用上的功能,但它指向一个更自然的未来:AI 不再等你把话说完。

📄 原文摘要(英文)

Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新