AI 搜资料,终于开始管「够不够全」而不是「像不像」
现在的 AI 搜索(比如 RAG、深度研究)会先抓一堆文档再挑重点,但挑的标准一直是「每篇单独像不像答案」,结果经常是十篇讲同一件事,真正缺的那块没人补。这篇把挑文档的标准从「单篇相关」改成「整组互补」:九个维度打分,再按整组表现反推每篇的贡献,冗余的扣分、关键的加分,不再让混子文档搭便车。训练时还按每轮结果的好坏给不同力度的提示,强的少管、弱的多教。在十个基准上,它整体最优,而且检索次数更少。它不是你明天就能用上的功能,但这是 AI 搜索从「找得像」走向「找得全」的一个真信号。
📄 原文摘要(英文)
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.