AI自己出题练自己,练着练着就废了
让AI自己出题、自己练,是现在提升推理能力的主流路子。但这篇发现一个反直觉的坑:练得越多,AI出的题越烂——要么是无效题,要么是同一道数学题换了个说法反复出现,最后模型性能不升反降。研究者拆出两个原因:无效题越练越多,而现有的查重只看字面相似,抓不住“本质相同、表述不同”的题。他们给的解法是让AI先学会识别并拒绝无效题,再用一个冻结的旧模型来给新题做“新颖度”打分,防止重复。结果在12个基准上稳定提升,连练十轮不崩,比现有方法高出17个点。这不是你明天能用的工具,但它解释了为什么“AI自己教自己”这条路会撞墙,以及怎么绕过去。
📄 原文摘要(英文)
Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filtering further increases their proportion in training data. Second, existing question diversity controls based on lexical similarity can miss mathematically equivalent questions expressed in different ways, which leads to question diversity collapse in later training rounds. Building on these findings, we introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. We first train the solver to recognize and reject invalid questions, then use its judgments to guide questioner rewards and filter solver training data. To avoid question repetition, we use a frozen base model to compare sampled question pairs and provide novelty feedback. Empirically, our method consistently achieves the highest average performance on 12 benchmarks in mathematical reasoning, general-domain reasoning, and code generation across two model families. Additionally, R-Quest maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.