AI Pulse
📄 论文解读

AI走进3D城市就迷路:局部看得懂,整体走不通

把大模型放进一座按真实香港数据建的3D城市里,让它以第一人称视角自己走、自己看、自己找路。结果很分裂:单看一张街景,它能认出地标、回答短距离的空间问题;一旦目标变远、路线变复杂,它就开始迷路,方向感尤其差,而且走错之后不会自己纠偏。研究者把这种失败叫「局部能力无法组合成持续的目标导向行为」——就像一个人每个单词都认识,但读不懂整段话。这不是某个模型的短板,而是当前多模态大模型的通病。它不是你明天能用上的东西,但它划了一条线:AI在真实城市里独立跑腿,还差得远。

📄 原文摘要(英文)

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新