让机器人自己改自己:不换大脑,只换工具
一个机器人模型,从头到尾没更新过参数,却能在任务中越做越好——靠的不是换大脑,而是让它自己给自己造工具、改流程。研究者把系统拆成两层:一个负责干活,一个负责看干活的结果、诊断哪里错了、然后修改工具和技能,甚至还能改进自己的诊断方法。这套自我改进的循环跑下来,机器人在模拟任务里的成功率从47%提到62%,在8个直接难倒原模型的任务上,成功率从1.25%飙到55%。更关键的是,这套在模拟里练出来的能力,搬到真实机械臂上继续自我修正,25次试验里84%成功。它不是你明天能买到的机器人,但它指向一个更省事的方向:与其不断训练更大的模型,不如让现有模型学会在行动中自我迭代。
📄 原文摘要(英文)
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.