GPT-6 Astra能当机器人身体吗?一半能,一半不能
GPT-6 Astra 不只是会规划,它还能直接输出机器人的动作数字。研究者把它接到真实机器人任务上试了六个领域,结果很分裂:导航它很强,指令跟随成功率 92%;但让它自己控制机械手去抓东西,手指协调就崩了。最扎眼的是速度——跑 30 秒的走路任务,模型要调用 250 次,平均每次等 40 秒,物理世界全程暂停等它思考。也就是说,它知道该做什么,但身体跟不上,脑子也太慢。这不是你明天能用上的东西,但它划出了一条线:大模型当机器人的“大脑”已经够格,当“小脑”还差得远。
📄 原文摘要(英文)
GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. In gripper manipulation, Astra can correct task targets and prepare contact conditions for subsequent policy execution; hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset. In dexterous manipulation, hybrid control achieves 50% success in ten experience-guided DexJoCo trials, while direct in-hand control struggles to coordinate finger contacts. In mobile manipulation, hybrid control reaches 38.7% success on the evaluated RoboCasa365. In navigation, Astra leads our local comparisons, reaching 92% success on RxR instruction following and 82% on HM3D object search, although search incurs substantial detours. In locomotion, dense motion-reference generation remains unreliable: none of five sequential attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. In humanoid loco-manipulation, Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks with pretrained whole-body controllers. These findings reveal a gap between useful task decisions and reliable physical control. Inference latency further constrains practical control: across 50 RoboDojo instances per condition, policy-assisted and direct control consume 624.8 million and 1.132 billion tokens. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference.