让机器人听懂人话去导航,不再靠死记硬背
现在的机器人导航,大多靠训练数据“喂”出来的:你问它没见过的任务,它就懵了。这篇的思路反过来——让大模型只负责“听懂你要什么、看懂周围环境”,真正怎么走路、怎么避开障碍,交给专门的导航工具去执行。机器人指哪走哪,还能从执行结果里自己修正路线,在没见过的场景里也能应付新任务。它不是你明天就能买回家的产品,但这是让服务机器人真正“通用”的关键一步。
📄 原文摘要(英文)
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments. Our key insight is to let the MLLM focus on interpreting requests, understanding scenes, and making decisions while preserving its general-purpose capabilities and delegating motion execution to navigation tools. To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM. Our harness supports these decisions with Navigation Skills, agent-oriented Tools for physical interaction, and task-progress and context management. A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback. Together, these components support sustained navigation across different task requirements and environments. SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks. Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments. Project Page: https://zju3dv.github.io/SuperNav/