AI Pulse
📄 论文解读

小模型也能当人类行为代理,大模型优势只在陌生任务

大语言模型被用来模拟人类行为,但研究者一直默认:模型越大越像人。这篇用 160 个心理学实验、1070 万条人类选择训练了 14 个不同规模的模型,结果反直觉:在训练见过的任务里,0.6B 到 1B 的小模型就能追平 70B 的大模型,规模几乎不重要。差距只在没见过的任务结构上拉开——大模型泛化更好。他们还做了个诊断:把任务说明、刺激、反馈、选择历史四类信息逐层剥掉,发现去掉刺激和反馈后模型表现暴跌 75.7%,说明它靠的是任务内容,不是靠猜历史选择。这不是你明天能用上的工具,但它意味着:心理学实验可以用小模型当“噪声下限”参照,省下大模型的算力,同时提醒我们——大模型在熟悉场景里的“聪明”可能只是记住了套路。

📄 原文摘要(英文)

Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新