AI Pulse
📄 论文解读

AI 特工终于能跨 App 干活了,靠的是把工具当乐高拼

现在的 AI 特工大多只会在一个 App 里打转:让它订机票,它不会顺手帮你查酒店、再同步到日历。这篇论文把「跨服务干活」变成了可批量训练的任务:先把一堆真实工具(API)封装成带标准接口的积木块,再让 AI 像写代码一样把这些积木按依赖关系随机拼成完整流程,自动生成并验证「信息要从 A 流到 B 再到 C」的任务。训练出的模型在 8 个基准上平均比原版强 9 分,在自动化测试里甚至超过了 Claude Opus 4.6。它不是你明天就能用的产品,但这是 AI 从「单机助手」走向「真正替你跑腿」的关键一步。

📄 原文摘要(英文)

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (CompoWorld), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新