AI 玩石头剪刀布,暴露了它不懂顺序
让 AI 玩石头剪刀布,结果暴露了一个尴尬的事实:它模仿的是你出拳的「频率」,不是「规律」。研究者让大模型在受控的两人对局里识别对手的隐藏策略、跟随简单的马尔可夫规则,结果发现:给更长的历史记录,它识别策略的能力并没有变好;就算它「认得出」规则,自己生成时也未必照着走;一旦规则涉及更复杂的条件依赖,它的表现大幅下滑。换句话说,AI 看起来行为像人,但背后的生成机制可能是错的——它只是学会了统计上的「大概这么出」,而不是真正理解「你上次出了布,所以我这次该出剪刀」。这不是你明天能用上的东西,但它提醒我们:当 AI 被用来模拟人类行为、做社会实验或游戏 NPC 时,表面上的逼真可能只是假象。
📄 原文摘要(英文)
Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.