AI Pulse
📄 论文解读

让AI写代码检查器,成功率不到一半

让AI修bug已经不够看了,这篇直接让AI从头写一个「静态分析检查器」——就是那种能自动找出代码里漏洞的工具。研究者从297个真实漏洞里造了300个任务,给AI一个带漏洞的代码库、一个修好的版本,让它自己琢磨出检查规则,再反复编译调试。21种模型配置跑下来,平均成功率32%,最强也就45%。也就是说,AI能修bug,但让它发明一套能抓同类bug的规则,还差得远。

📄 原文摘要(英文)

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新