AI Pulse
📄 论文解读

小模型想破头不如问一句

我们总以为 AI 多想想就能变聪明;这篇发现,小模型想再多,也只是在已知答案里打转,真正缺的是它不知道的知识。研究者拆开推理过程,把失败分成两类:一类是路走岔了,回头能找回来;另一类是知识不够,想破头也没用。于是他们训练了一个 4B 和 8B 的小模型,让它先自己推理,判断卡在哪,如果是知识不够,就去问更大的模型。结果,4B 的小模型在 1158 道难题上超过了 14B 的大模型,服务成本还低了 2.7 倍。这不是你明天能用的工具,但它指出了一个方向:让不同大小的模型分工,而不是一味堆算力。

📄 原文摘要(英文)

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新