AI 规划失败,可能不是算得不准,而是它衡量距离的尺子歪了
AI 在脑子里做规划时,靠的是把目标状态和当前状态在「想象空间」里比距离,距离近就认为更可行。但研究者发现,训练时为了让模型不崩溃而加的常规约束,会让这把尺子量出来的远近和真实任务的难易对不上——它可能把一条根本走不通的路判成捷径。这篇提出一个改动:不再用固定形状的球来衡量距离,而是让模型自己学一个可变的椭圆尺子,只调整训练方式、不碰推理逻辑,在四个视觉控制环境里规划成功率全部提升。它不是你明天能用上的东西,但它点破了一个容易被忽略的真相:AI 的「直觉」好不好,往往取决于它默认用哪把尺子量世界。
📄 原文摘要(英文)
Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with ΛReg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/