AI Pulse
📄 论文解读

AI 智能体也能进化了:模型不动,靠「自然选择」变强

大模型智能体的能力不只取决于模型本身,还取决于它的「装备」——提示词、工具、技能、控制流程。以前让 AI 自我改进,就像一个人单打独斗地试错,改好一个任务常常搞砸另一个。这篇论文换了个思路:把「装备」当成一个种群,让它们像生物一样自然选择——只保留那些不倒退、还能扩展能力的变体,同时保留其他分支作为后备,失败、老师、自己得出的经验都走同一个修改接口。结果在四个基准测试上平均提升约 17 分:终端操作从 75.5% 涨到 83.2%,真实网页任务从 43.5% 飙到 93.0%,而且学到的能力能直接迁移到没见过的任务上。关键点:模型权重完全冻结,变强的是「装备」本身。这不是你明天就能用的工具,但它揭示了一个趋势——AI 的进化可能不再需要换大脑,而是靠优化「生存装备」。

📄 原文摘要(英文)

An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新