AI自己给自己写游戏,还能越改越好
让AI自动生成一个能玩的游戏不难,难的是让它自己把游戏改到好玩、没bug。这篇提出一个叫RSIGame的框架,让AI像程序员一样干活:先自己玩一遍游戏,找出问题,修完再玩,循环往复;同时记一本“问题清单”,把踩过的坑和修法都记下来,下次生成新游戏直接调用。在140个任务、两个游戏引擎上,它稳定超过普通的一次性生成。最狠的是,一个27B的小模型用这套方法,生成质量超过了GPT-5.5的一次性结果,而且生成成本还低了11倍。它不是你明天就能拿来开发游戏的东西,但“让AI自己当自己的测试员和修理工”这个思路,正在把AI从“生成一次就完事”推向“持续自我改进”。
📄 原文摘要(英文)
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.