AI Pulse
📄 论文解读

AI终于能干活干到底了:一个35B小模型干翻大模型

现在的AI聊天很聪明,但一让它干点正经活——比如整理文件、跑代码、查资料再写报告——就经常半路卡住、忘掉上下文、或者干脆编个结果糊弄你。这篇论文把这个问题叫「工作能力」:不是会聊天,而是能持续、可验证地推进一个真实目标。

他们做了两件事:一是让AI能操作更多真实环境(文件、搜索、代码),并且每一步都能被验证;二是训练AI学会拆解长任务、并行干活、等结果回来再整合、发现不对就重来。核心是让AI自己维护一个「任务状态」,像项目经理一样跟踪进度和来源。

结果很有意思:一个只有350亿参数的模型(Apodex 1.1 Mini),在金融、科研、编程等复杂任务上,干翻了很多参数大它几倍的模型。这意味着你不需要租天价算力,就能在本地跑一个真正能帮你干活的AI。

它不是你明天就能装上的工具,但方向很明确:AI正在从「聊天玩具」变成「能交差的下属」。

📄 原文摘要(英文)

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this working capability: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. Environment Scaling expands the diversity and verifiability of executable file, search, and code environments, while Agentic Coordination Scaling trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a Heavy-Duty Solver for ambitious, long-running tasks.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新