AI Pulse
📄 论文解读

AI 终于能同时看见两个玩家的世界

过去让 AI 预测第一人称视角,一次只能管一个人;这篇让 AI 同时生成两个玩家的第一人称画面,而且两人在同一个世界里互动——你推门,我这边门也动了。做法是把两个视角的画面放进同一条生成序列里一起处理,再给每个视角喂上对方的方位和共享的场景记忆,保证两边看到的世界对得上。它不是你明天能用上的东西,但这是 AI 从「单机游戏」走向「联机游戏」的关键一步:多智能体协作、具身交互的模拟,从此有了更真实的基础。

📄 原文摘要(英文)

Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新