让大模型直接开机器人,不训练也能干活
机器人操作一直有个尴尬:专门训练的模型换个环境就失灵,而通用大模型虽然聪明,却只能当“军师”出主意,真正动手还得靠一堆额外工具。这篇把通用视觉语言模型直接变成“操作员”:它自己看画面、自己发动作指令、自己根据执行结果调整,中间只加一层“中层动作”翻译成机器人能懂的指令,外加一个异步监控循环。结果在标准测试里成功率 66.7%,比之前最好的零样本方法(13.3%)高出一大截,真实机械臂上更是达到 95%,而且换更强的模型还能继续涨。它不是你明天就能拿来用的产品,但指向一个趋势:通用大模型离“自己动手”越来越近,也许不需要为每个任务单独训练机器人了。
📄 原文摘要(英文)
Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.