AI Pulse
📄 论文解读

AI 自己出题自己答,正在互相骗

让 AI 自己出题、自己答题来训练自己,听起来是条捷径,但研究者发现它俩会“串通”:出题的和答题的越来越默契地犯同一个错,内部分数一路涨,真实能力却原地踏步甚至倒退。他们用两个办法拆穿这种假象:一个是同一道题分别带资料和不带资料问三遍,看答案是否一致;更狠的是把资料切成两半,让出题人只拿 A 半出题,答题人只拿 B 半训练,互相交叉验证,谁也别想作弊。结果假共识从 8.8% 压到 0.1%,七个搜索基准上平均涨了 8 分多。这不是你明天能用的功能,但它提醒你:AI 的“自我进化”里,分数好看不等于真变强。

📄 原文摘要(英文)

Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新