AI Pulse
📄 论文解读

AI 搜索进化:让模型自己决定何时查资料

进化算法让大模型自己改答案,但卡住时它缺外部知识,硬塞搜索又可能反复翻同一页。EvoDuet 给模型装了个「检索门」:每步先自评知识缺口,决定是查新资料、复用旧资料还是直接继续;内层循环按「这份资料预计能带来多少分数提升」来排序,外层循环再并行生成候选。在 21 个优化任务上,它把 OpenEvolve 的发现增益从 74.1% 提到 78.0%(GPT-5.6-Luna)、从 61.3% 提到 82.3%(Gemini-3.8-Flash),并在 8 个任务上刷新了此前最优成绩。这不是你明天能直接用的工具,但它指向一个趋势:AI 不再被动等喂资料,而是学会主动判断「我缺什么、该去哪找」。

📄 原文摘要(英文)

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新