AI Pulse
📄 论文解读

手机AI助手一遇弹窗就崩,最强模型也扛不住

你让手机上的AI帮你订机票、改设置,它正干着活,突然弹出一个权限请求或系统提示——大多数AI当场就懵了。研究者造了个叫AnTrap的测试场,往AI干活的过程中随机扔各种意外:弹窗、点错、界面卡死,结果16个主流模型全军覆没,最强的那个也掉链子。更关键的是,他们发现有些坑AI能靠多训练学会躲,但有些坑——比如界面卡死这种需要真正理解上下文的——怎么练都白搭。这不是你明天就能用上的功能,但它告诉你一件事:现在的手机AI远没到能放心交给它办事的程度,它只会在风平浪静时显得聪明。

📄 原文摘要(英文)

GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies. We introduce AnTrap, a comprehensive benchmark that injects dynamic perturbations into agent execution trajectories. We propose a taxonomy organizing real-world anomalies into four layers (State, Thinking, Action and Round) with ten fine-grained subcategories, and develop a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models, we reveal universal vulnerability to dynamic anomalies, with even the strongest models suffering significant performance degradation. Furthermore, we conduct GRPO training in both original and adversarial environments to validate our benchmark, separating environment-learnable anomalies from reasoning-bottlenecked ones. Our findings show that while single-step traps at state and action layers are largely addressable through adversarial reinforcement learning, deep contextual traps, like state deadlock, expose intrinsic limitations that cannot be resolved by training in environments with traps alone.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新