AI Pulse
📄 论文解读

教机器人新技能,不再需要几百小时人工录像

现在教机器人一个新动作,主流办法是让人远程操作它几百上千次,录下数据再训练,成本极高。这篇论文换了个思路:让机器人自己“想”着练。它先让一个会“看图说话”的模型把复杂任务拆成几步,机器人照着拆好的步骤自己试、自己收集数据,再模仿自己的成功经验,最后用强化学习微调。结果是用远少于以前的演示数据,就能学会新任务,效果还更好。它不是你明天就能用上的东西,但“让机器人自己练”这条路,可能让未来教机器人新技能像教人一样自然。

📄 原文摘要(英文)

How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular, collecting teleoperation datasets requires hundreds of hours of expensive human labour and the alternative, reinforcement learning (RL), can be notoriously sample-inefficient especially for long-horizon tasks. In addition, RL with VLAs imposes several challenges due to the model's size and architectural design. In this work, we propose EXIMO, an efficient algorithm for finetuning of VLA policies. EXIMO operates in three stages: explore, imitate, and optimize. During the explore phase, EXIMO equips the VLA with a vision language model (VLM) that acts as a planner. The VLM thinks and breaks down challenging long-horizon problems into shorter ones for the VLA. The VLM, together with the VLA, is used to collect an orchestrated dataset on new tasks. During the imitate phase, the VLA is finetuned with the orchestrated data. Finally, during the optimize stage, we use residual off-policy RL to further finetune the policy. In our experiments, we ablate all three stages of EXIMO and show that it outperforms existing approaches significantly in terms of sample-efficiency and final performance.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新