AI角色扮演评测翻车:固定剧本测不出真实体验
现在的AI角色扮演评测有个大坑:它让AI接着一段固定对话往下演,再用一套固定标准打分。但真实用户不是这么玩的——每个人口味不同,对话走向也千变万化。这篇论文干脆让AI模拟真人用户,跟角色扮演AI自由聊天,再给每个模拟用户定制专属评分标准。结果发现,定制评分比通用评分更贴近真人感受。它用300个角色档案、5个模拟用户,测了16个AI,能分别看出对话质量、长线能力和个人体验。这不是你明天能用的工具,但它点破了评测界的通病:拿固定剧本考AI,考不出它跟真人聊天的真实水平。
📄 原文摘要(英文)
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.