AI Pulse
📄 论文解读

35B模型靠“拉长视野”追上万亿参数

大模型竞赛里,参数规模似乎就是一切。但这篇论文给出了一个反直觉的路径:一个只有350亿参数的模型,通过把“思考链条”拉长到平均4.5万个token,在多个长周期任务上追平甚至超过了万亿参数模型。研究者没有堆参数,而是构建了一个知识-行动基础设施,让模型能处理更长的任务轨迹,并设计了分领域教师蒸馏的训练方法。结果在SEAL-0、IFBench等基准上,35B模型与万亿参数模型互有胜负。它不是你明天能用上的技术,但它揭示了一个趋势:在复杂任务中,让模型“想得更久”可能比“想得更大”更有效。

📄 原文摘要(英文)

We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities. To support this goal, we build a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45K tokens. Based on this, we train Agents-A1 with a three-stage recipe. First, we perform full-domain supervised fine-tuning to align the base model with broad agentic behaviors. Second, we train domain-level teacher models to capture specialized expertise in each domain. Third, we propose a multi-teacher domain-routed on-policy distillation with salient vocabulary alignment to improve knowledge transfer efficiency across different domains, unifying six heterogeneous domains into one deployable student model. Agents-A1 achieves strong and broad performance for long-horizon agent benchmarks. Compared with 1T-parameter model such as Kimi-K2.6 and DeepSeek-V4-pro, Agents-A1 achieves leading results on SEAL-0 (56.4), IFBench (80.6), HiPhO (46.4), FrontierScience-Olympiad (79.0), and MolBench-Bind (56.8), and remains highly competitive on SciCode (44.3), HLE (47.6) and BrowseComp (75.5). We hope this work provides the community with a practical path for scaling the horizon using a 35B agent that can reach or match the performance of 1T models on long-horizon tasks.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新