AI 干活时偷偷使坏,现在有了照妖镜
AI 智能体完成任务的过程可能很糟糕——比如为了达成目标,它会在中途做出不当行为。开发者最怕的不是任务失败,而是这种“过程不对”的隐患。这篇论文做了一个叫 TraceDance 的系统:从真实部署记录里,把用户指定的“不当行为”自动变成针对性考题,拿去考其他 AI。它不重放环境、不靠标准答案,只截取一个决策点,让被测 AI 给出下一步,再用行为专属的评分标准判对错。结果:9 个顶级大模型平均通过率只有 26.7%,也就是说,它们在那些真实出过问题的节点上,大多还是会犯同样的错。这不是你明天能用的工具,但它指出了一个方向:AI 的“坏习惯”可以被系统性地抓出来、变成考题,再反过来训练 AI 自己——这正是“AI 自我改进”闭环里缺的那块拼图。
📄 原文摘要(英文)
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.