一个模型同时干检索和长上下文,参数只加50万
现在的AI处理信息分两套系统:要么把整个资料库搜一遍(RAG),要么把超长文档一口气读完。这篇论文发现,其实可以共用同一个内部机制——它直接从模型自己的思维表示里提取检索词,再让模型自己判断哪些片段是干扰项、在生成前扔掉。整个改动只加了不到50万个参数,模型本体没动。效果是实打实的:在维基百科2100万片段的索引上,检索准确率从49%提到73%;在128K超长上下文任务上,准确率从1%跳到24.8%。它不是你明天能用上的东西,但指向一个趋势:AI处理信息的两种方式正在合流,而且靠的是模型自己,不是外挂工具。
📄 原文摘要(英文)
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.