AI Pulse
📄 论文解读

AI角色扮演评测翻车:固定剧本测不出真实体验

现在的AI角色扮演评测有个大坑:它让AI接着一段固定对话往下演,再用一套固定标准打分。但真实用户不是这么玩的——每个人聊法不同,感受也不同。这篇论文干脆让AI模拟出5个不同性格的“虚拟用户”,跟候选AI自由对话,再给每个用户定制打分标准。结果发现,定制标准比通用标准更贴近真人判断。它把“AI角色扮演好不好”从单一排名拆成了“对谁、聊多久、体验如何”的细账。这不是你明天能直接用的功能,但它意味着以后选AI伴侣或情感陪伴产品时,评测会更能反映“你”的真实体验,而不是一个平均分。

📄 原文摘要(英文)

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新