AI 训练后变聪明,但代价是失去更多可能性
一个反直觉的发现:大模型经过强化学习训练后,单次回答的准确率确实更高了,但如果你给它足够多的尝试机会,它反而可能比训练前解决更少的问题。研究者把这种代价叫做「锐化税」——训练把任务推向两个极端:要么总能答对,要么永远答错,于是模型变得稳定、高效,却牺牲了探索的广度。这个现象在需要多轮调用工具、与环境交互的「智能体任务」里同样普遍,横跨 14 组模型、42 个测试场景。好消息是,他们提出了一种简单的采样方法,能在训练时少交这笔税,既提高单次准确率,又保留多次尝试时的覆盖面。它不是你明天就能用上的东西,但它提醒我们:AI 变强的方式,可能正在悄悄收窄它解决问题的可能性空间。
📄 原文摘要(英文)
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.