AI 写代码只会改一个仓库,跨仓库协作直接抓瞎
现在的 AI 编程助手,评测都只在一个代码仓库里考它。但真实软件生态里,一个功能常常要同时改好几个仓库——比如改完主程序,还得同步改依赖库和测试。这篇论文专门造了 120 个跨仓库的真实任务,结果最强配置成功率也只有 42.5%,多数 AI 要么压根没发现要改别的仓库,要么改了但没改完。更关键的是,让 AI 同时看多个仓库比一个个来效果好得多——它能从相关仓库里找到线索,知道自己改得对不对。这不是你明天能用的功能,但它划出了一条真实的分界线:AI 写代码的“单打独斗”已经不错,可一旦需要跨项目协作,它离人类工程师还差得远。
📄 原文摘要(英文)
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.