AI Pulse
📄 论文解读

AI 操作电脑的新思路:图形界面太慢,命令行来提速

现在的 AI 操作电脑,要么只会点鼠标(图形界面),慢还容易点错;要么得给每个软件单独写接口,工程量大到没法推广。这篇论文提出一个中间路线:让 AI 自己决定什么时候点鼠标、什么时候敲命令行——比如复制文件、查系统信息这类事,敲命令比点几十下快得多。他们造了 8 千条训练数据,教模型在两种操作间自由切换,再用强化学习奖励它“该用命令时就用”。结果在 OSWorld 测试上准确率从 38.8% 提到 53.6%,在另一个 Windows 平台上也涨了 4 个点,说明这套思路能跨系统用。它不是你明天就能用上的产品,但这是 AI 从“会看屏幕”走向“真会干活”的关键一步:未来的 AI 助手可能不再只是替你点按钮,而是像熟练工程师一样,该点就点、该敲就敲。

📄 原文摘要(英文)

Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新