机器人失败后怎么补救,现在能写成经验手册传给下一个AI
机器人干活失败时,通常只能重试或卡住。这篇让一个“强AI”先试错,把失败后怎么调整动作写成一本“操作手册”,再交给一个“轻AI”照着执行;轻AI执行时发现手册不管用,强AI再改,如此反复,手册越磨越准。最后这本手册不用改模型参数,直接给任何机器人用:真实操作成功率从37.3%提到64%,轻AI带上手册甚至超过没带手册的强AI。它不是你明天能装进自家机器人的东西,但指向一个方向——机器人的经验不再只锁在训练数据里,而是能像老师傅的笔记一样传下去。
📄 原文摘要(英文)
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.