AI Pulse
📄 论文解读

AI 会写代码,但一碰真实世界就露馅

现在的 AI 能写代码、能操作电脑,但把它扔进物理世界会怎样?研究者造了个 84 项任务的模拟考场:机械臂抓取、移动、开车、飞行,全都要 AI 自己看画面、做判断、动手操作。结果发现,AI 能完成很复杂的子任务——比如图像分割、相机标定、估算空间位置——但一组合起来就崩:明明走到了目标位置,却把要抓的东西跟丢了;动作没效果也不改;任务没做完就宣布完成。不同模型还各有短板,一个擅长空间定位,另一个擅长保持平衡。它不是你明天能用上的东西,但这份考卷的价值在于:它把「AI 离物理世界还有多远」从感觉变成了可量化的差距清单。

📄 原文摘要(英文)

General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新