机器人看一遍演示就会做,不再需要海量训练
机器人学新任务,过去要么靠人写死程序,要么靠海量数据训练。这篇把「看一遍演示就会做」这件事重新定义清楚:一段演示视频里其实混着动作轨迹、物体语义、空间关系、任务目标,机器人到底该学哪个,以前没人说清。研究者给出明确的学习目标,并搭了一套极简框架——一个视觉提示编码器加低成本数据采集,不靠大规模预训练,在仿真和真实环境都表现不错。它不是你明天就能用的产品,但它把「机器人看演示学任务」从玄学变成了可复现的工程,意味着这个方向的门槛被拉低了。
📄 原文摘要(英文)
We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including action, semantic, composition, and affordance discrimination. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.