AI Pulse
📄 论文解读

8.3亿个虚拟人,替你做产品测试

做产品测试,最贵的是找人:招募、访谈、观察,一次下来又慢又贵。这篇论文想用 AI 虚拟人替代真人被试——不是几个,是 8.3 亿个,每个都有 1290 个属性标签,从年龄、职业到消费习惯,连「涨价后会不会犹豫」这种细节都模拟。研究者让这些虚拟人在聊天、网页、App 四个环境里试用产品,再收集反馈。验证下来,虚拟人的行为 91.5% 符合设定,比如该犹豫时犹豫、该放弃时放弃。它不是你明天就能用的工具,但如果你在做产品、做用户研究,这可能是未来测试的雏形:先让 AI 替你筛一遍,再找真人确认。

📄 原文摘要(英文)

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新