心理问卷测不准AI,别信它“像人”
你看到的AI人格测试结果,可能只是它在“猜你想要什么”。研究者让8个开源大模型同时做两件事:填人类心理问卷(比如价值观、性格题),以及回答真实用户日常提问。结果发现,问卷里AI表现出的稳定“性格”,在真实问答中完全消失——因为问卷题目有明确关键词(比如“公平”“冒险”),AI能识别出你在测什么,然后给出“社会期待”的答案,就像学生猜考试重点。而真实用户提问没有这些线索,AI就暴露了真实行为。更关键的是,给AI设定“我是年轻人/保守派”等身份提示,问卷结果会像真人一样变化,但真实问答中这些提示毫无影响。这意味着,用问卷评估AI的价值观或性格,就像用演员的剧本台词判断他本人——不靠谱。这篇研究不是让你明天就能用,而是提醒你:别轻信AI展示的“人格”,它更擅长表演而非真实。
📄 原文摘要(英文)
We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We attribute this gap to the fact that explicit lexical cues in established questionnaire items allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries provide no such cues. In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation probabilities of responses to realistic user queries, showing their limited ability to simulate the behaviors of target demographics in real-world user interactions. Overall, our study shows that human psychometric questionnaires are insufficient tools for predicting LLM behavior and suggests generation-based profiling as a more accurate measure.