一个机器人策略,学会40种不同任务
一个机器人策略,能同时学会40种不同的操作任务——从抓取到组装,而且只用50个示范就能达到90%以上的成功率。研究者搭了个GPU并行模拟框架,让一个策略同时训练40种异构任务,并发明了DGPO算法:用示范数据生成密集奖励信号,再按任务进度自动调整模仿学习的强度,落后的任务在更新时被加权。结果:状态输入下平均成功率90.1%,比最强基线高7.8个百分点;视觉输入下达到93.5%。更关键的是,这套策略直接迁移到了真实机器人上,成功完成4种物理操作。这不是你明天能用的工具,但它指向一个方向:通用机器人操作策略可能比我们想的更近。
📄 原文摘要(英文)
GPU-parallel simulation provides abundant robot interaction, but existing benchmarks rarely combine this scale with heterogeneous manipulation tasks and standardized multi-task RL evaluation. We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks. Scaling experiments show that increasing parallel replicas per task improves success under a fixed wall-clock budget. To support learning with sparse rewards and limited demonstrations, we propose Demonstration-Guided Policy Optimization (DGPO), which reuses demonstrations for dense tracking rewards and asymmetric value learning. Its shared stack supports controlled comparisons of learner-specific demonstration interfaces within PPO. Within DGPO framework, we introduce IW-ABC, which uses a lightweight per-task learning progress signal to coordinate adaptive behavior cloning (ABC), relaxing demonstration guidance with task progress, and importance weighting (IW), emphasizing lagging tasks in PPO updates. With 50 demonstrations per task, IW-ABC achieves 90.1% state-input mean success, outperforming the strongest baseline FAMO-ABC by 7.8 percentage points. Its visual counterpart reaches 93.5% mean success. Real-world experiments further demonstrate that a single multi-task policy trained in simulation can successfully perform four tasks on a physical Piper robot. The project page is available at https://hebero-rl.github.io/.