AI Pulse
📄 论文解读

给AI装个进度条:它终于知道自己做到哪了

机器人干活越来越久,但AI一直有个盲区:它不知道自己做到哪一步了。你让它叠衣服,它可能叠到一半就懵了——因为只看当前画面,根本看不出这是第几件。这篇论文发现,连最强的进度评估模型也会迷路,但只要给它补上“之前发生了什么”的上下文,同样的模型错误率直接砍掉近八成。研究者还做了个自动补上下文的循环,让AI自己去找线索,进度判断准确率提升63%。这不是你明天能用的功能,但它解释了为什么机器人总在长任务上掉链子,也指了个明确的方向:让AI学会回头看。

📄 原文摘要(英文)

Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新