AI Pulse
📄 论文解读

机器人没学过的新任务,靠“现场进化”自己搞定

一个机器人控制器没训练过的新任务,通常意味着要重新训练。这篇让它在测试时“现场进化”:不重训,只靠重新组合自己已有的技能,从自己的失败里改进,并且记住学到的东西。做法是把任务写成“奖励程序”——分阶段、带完成条件和可调参数;一个大语言模型根据执行反馈不断改写程序结构,数值优化器调参数,每个候选都在并行模拟里验证。结果是人手设计的奖励只挖掘了模型一小部分能力,而进化出的程序能释放更多,有时还走出人类没想到的策略,并且真的跑在实体 Unitree G1 机器人上。它不是你明天能用上的东西,但指向一个方向:机器人不再为每个新任务重训,而是像人一样,拿已有本事临场发挥。

📄 原文摘要(英文)

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新