AI 写游戏能开局,但修 bug 和防回退还不行
让 AI 从一句话做出一个能玩的游戏,这事已经不算新闻了。真正难的是后面:游戏坏了它能不能自己找到毛病、修完别把别处弄坏、再按你的要求一版版改下去。研究者把整个开发过程拆成三个阶段——生成、修 bug、多轮优化,做了 97 个生成任务、100 个修 bug 任务(每个游戏里埋了 19 到 27 个错)、17 条六轮优化链,用真人试玩和行为测试来打分。结论很直接:现在的 AI 搭出能跑的游戏、照着明确要求做,都挺稳;但让它自己发现缺陷、验证运行时行为、改完不破坏原有功能,明显拉胯。这不是你明天能拿来当游戏开发工具的东西,但它划出了一条真实的能力分界线:AI 当执行者可以,当质检员和守门员还早。
📄 原文摘要(英文)
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.