多智能体撞车难题:先商量再行动,成功率逼近满分
多智能体路径规划里,每个机器人各自选路,单看都对,合在一起却会撞车——问题出在“最后一步各自拍板”。这篇把“一次定死”改成“先随机起个念头,再和邻居来回对几轮,逐步收敛到不冲突的联合行动”,像扩散模型去噪一样。在1600个任务上,微调后解出1598个,成功率碾压现有方法,还能扩展到百万级智能体。它不是你明天能用上的东西,但“先沟通再承诺”这个思路,对任何多智能体协作系统都是值得盯着的方向。
📄 原文摘要(英文)
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same context, independently sampling from per-agent distributions can recombine locally valid choices into incompatible joint actions. This failure can arise from the final sampling mechanism even when the per-agent action distributions are learned correctly. DMM (Decentralized Master-Mind) addresses this by replacing one-shot action sampling with discrete, iterative refinement of action intents across communication rounds, inspired by denoising in diffusion models. Agents initialize random action intents and refine them through local communication, coupling their choices before commitment. DMM is pretrained with imitation learning on expert MAPF solutions and further optimized with MICPO, a critic-free group-relative reinforcement-learning method designed for multi-agent, multi-round action refinement. DMM generally achieves higher success rates and lower solution costs than the evaluated learnable baselines. On 1,600 MovingAI tasks, DMM fine-tuned with MICPO solves 1,598, the highest coverage among the evaluated methods, while achieving solution costs close to those of the strongest baselines. DMM also scales to over one million simultaneously acting agents in obstacle-rich environments. These results show that round-level intent refinement can improve joint-action coordination while preserving decentralized execution.