AI看视频学技能,卡在“该不该学新东西”
让机器人看视频学动作,研究者发现一个反直觉的瓶颈:模型很会认出“这是在做同一件事”,却几乎不会主动说“这个我不会,得新学一个”。他们让19个开源视觉语言模型看机器人桌面操作和人类下厨的视频,要求把连续画面切成动作、把相同动作归组、并决定是复用已有技能还是新建一个。结果:很多模型连“把相同动作归组”都接近瞎猜,规模变大也不一定变好;更关键的是,即便经过专门训练,模型能巩固见过的技能,但技能库永远长不大——训练里没见过的动作,它能定位到时间点,却几乎从不为它新建技能。这暴露了具身智能的一个深层短板:识别“自己不会”比“会做”难得多。它不是你明天能用上的东西,但指出了让机器人从视频里自学动作这条路真正的拦路石在哪。
📄 原文摘要(英文)
Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.