AI Pulse
📄 论文解读

AI 学操作电脑,不再需要真电脑

训练 AI 操作电脑,过去得先给它看大量真人操作录像,而录像只能来自真实软件——装系统、开应用、点按钮,成本高、场景有限。这篇论文换了个思路:让 AI 自己“演”操作过程。它先用文字描述一个桌面长什么样(系统、窗口、按钮位置),再用图像生成器把画面画出来,然后让一个规划器决定“下一步该点什么、屏幕会变成什么样”,图像生成器再画出点击后的新画面。就这样,AI 在想象中完成了整个操作流程,全程没碰过真实软件。用这套方法生成的 7.9 万条训练数据,让 AI 在真实桌面任务上的成功率从 33% 提到 40.8%,在科学软件任务上从 14% 提到 32.2%。它不是你明天能用上的功能,但它意味着:AI 学操作电脑的门槛,从“必须有一台装满软件的电脑”降到了“只要会画图”。

📄 原文摘要(英文)

GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新