AI Pulse
📄 论文解读

让AI学会“临时抱佛脚”:不靠背答案,靠读文档

大模型最擅长的是“背答案”——训练时见过的知识它张口就来,但给它一篇新文档、让它现学现用,它就露怯了。这篇论文的做法很反直觉:故意把公开文档“弄脏”一点,再让AI自己出题、自己答题,只保留那些必须依赖文档才能答对的样本,用这些数据训练学生模型。结果,一个350亿参数的模型在“从上下文学习”的基准上,从13.7%涨到24.6%,超过了万亿参数的顶级模型。它不是你明天能用上的功能,但它指出了一个方向:AI的“临场学习”能力,可能不需要昂贵的人工标注,靠“折腾”现成文档就能练出来。

📄 原文摘要(英文)

Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新