AI 自己给自己写规则书,然后照着升级
大模型通常被冻住:训练完就定型,新环境里只能靠提示词临时应付。这篇让模型把对世界的猜测写成一本「规则书」存进外部记忆,再把它编译成可执行代码,靠预测错误不断修订规则、重写代码,而且每次更新都要通过「重放验证」——代码跑出来的结果必须和真实观察一模一样才被采纳。在 ARC-AGI-3 的 25 个公开游戏上,它全部通关,平均人类操作效率 100%,只用了人类 44% 的操作次数;在 Atari 乒乓球上,学出的控制器三局全胜 21:0,全程不再调用大模型。它不是你明天能用上的东西,但指向一个更吓人的方向:模型不再靠人喂新数据变强,而是自己探索、自己改自己的认知、自己验证——底层模型没动,能力却在长。
📄 原文摘要(英文)
Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.