AI Pulse
📄 论文解读

AI 干活时总漏东西,给它加个便签

AI 帮你从一堆文件里做报告,它读过的、没读的、任务要求的,全都挤在上下文窗口里,转头就忘——经常图都抽出来了,报告里却没放。这篇给 AI 配了个「环境侧便签」:任务要求记在便签上,每读一个文件就记一笔出处,列了没打开的文件也标成候选;AI 能对着便签逐条核对,没做完想交差,系统会拦下来提醒。三个模型、三个基准上都比裸干强。它不是你明天就能用的产品,但「让环境替模型记状态」这个思路,是 AI 干活靠谱化的一个真方向。

📄 原文摘要(英文)

Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新