AI 智能体在复杂环境里更容易被攻破
现在的 AI 安全测试大多只考短对话,但真正的 AI 智能体要在长期任务里反复改状态、做几百步操作,风险是累积的。这篇论文造了一个 1 万多个场景的测试场,让 AI 智能体平均要调用 97 次工具才能完成任务,再用一种黑盒攻击方法去不断改变环境状态、逼它出错。结果:任务越复杂,攻击成功率越高,从简单场景的 2% 优势一路拉到复杂场景的 17% 以上。更值得注意的是,同一个模型,跑在不同框架里,安全表现差异很大——也就是说,安全漏洞不只在模型脑子里,也在工程实现里。它不是你明天能用上的工具,但给「AI 智能体到底哪里会翻车」画了一张更真实的地图。
📄 原文摘要(英文)
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.