AI Pulse
📄 论文解读

AI 干活干到一半会“卡住”,这篇教它先想清楚自己缺什么

大模型做长任务时,不是记不住,而是不知道自己“卡住了”。研究者给 AI 加了一个显式的“信念状态”:它每走一步都先写下两件事——现在世界是什么样、任务还差什么。然后它自己检查这个信念是否自洽、有没有在空转;一旦发现“信念陷阱”(一直在动、但没靠近目标),就按卡住的方式和缺的东西分别对症恢复。四个基准、三种模型上都是最高分。它不是你明天能用上的,但指向一个方向:AI 的“记忆”不该只是存历史,而该是不断更新的世界观。

📄 原文摘要(英文)

Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新