AI Pulse
📄 论文解读

机器人不再需要AI大脑:纯代码决策,成功率70%

机器人控制一直默认要有个大模型在回路里:视觉-语言-动作模型看画面出动作,或者运行时再问一次视觉大模型。这篇论文提出一个反直觉的视角:把机器人和环境的状态当成一台图灵机的纸带,规则就是代码。只要状态量得准,决策可以完全用代码写出来,不需要任何模型在线推理。研究者用相机图像和本体感觉实时追踪状态,所有判断都从代码里出,跨任务共享一个库。在42个双臂操作任务上,这套纯代码方案测试时零模型,成功率70.24%。好处很直接:状态显式可查、执行可控可恢复、新任务能复用旧代码,能力可以累积。这正好是递归自我改进的载体——让写代码的AI自己迭代这个库,每次改动都看得见、控得住。它不是你明天能用上的产品,但它挑战了一个默认假设:也许机器人不需要更聪明的模型,只需要更准的状态描述和更稳的逻辑。

📄 原文摘要(英文)

Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新