AI 自己检查自己:分歧暴露真相,共识反而藏错
让 AI 干长活,最难的不是干活,是验收。这篇论文发现一个反直觉的现象:让同一个模型多跑几遍,结果不一致的地方往往藏着正确答案,而大家一致的地方反而可能一起错。于是研究者把模型本身改造成一个「验收员」:给它工作台、查证工具和可复用的检查技能,让它专门去挑毛病——分歧处查证据,共识处找漏洞。在五个长任务基准和两个前沿模型上,这套方法选出的最终答案比单次生成平均高 6 分左右,而且检查技能还能从失败中自我改进。它不是你明天能直接用的产品,但它指向一个趋势:AI 的自我纠错,正在从「让模型重写」升级为「让模型当质检员」。
📄 原文摘要(英文)
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.