AI 在陌生环境里学不会,除非把每一步都记下来
我们总以为 AI 靠预训练知识就能应对新环境;这篇用全新设计的文字游戏证明,它连规则反着来的游戏都玩不转,除非把每次尝试的完整记录都留着,而不是总结成规则。研究者造了一批规则新颖甚至反直觉的游戏,让 AI 只能靠反复试错学习,结果发现:保留完整行动和反馈记录比提炼成策略更有效;人类顶尖玩家比最强 AI 得分更高,因为人类尝试更多样、重复更少;换掉 AI 的「外壳」框架,在同样大脑下能提性能还降成本。它不是你明天能用上的东西,但提醒你:AI 的「经验」不是总结出来的,是记下来的。
📄 原文摘要(英文)
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing benchmarks have sought to evaluate this ability, but they primarily evaluate tasks whose rules are provided in the instructions or already familiar to pretrained models, making it difficult to distinguish learning from interactions from reasoning with existing knowledge. To address this, we introduce Learn2Play Bench, a benchmark of newly designed text-based games, whose rules are novel or counterintuitive, requiring agents to acquire knowledge through interaction rather than rely solely on pretrained knowledge. These games provide reproducible feedback and automatic scoring, enabling controlled evaluation of learning across repeated attempts. We also vary game instances to test whether agents can apply what they have learned to new situations. Therefore, we evaluate how backbone models, self-evolving methods, and agent harnesses affect agents' learning ability, revealing three findings: (1) Experience retention: Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies. (2) Human agent gap: Top-performing human players achieve higher peak scores than the evaluated agents. Human explore more varied strategies, and repeat actions less. (3) Harness matters: With the backbone fixed, changing the harness can improve performance while reducing estimated inference cost. Together, these findings provide insights into how LLM agents learn from experience and suggest directions for future work to improve their learning ability. Project website: https://liushiliushi.github.io/learn2play-bench-website/