AI Pulse
📄 论文解读

AI 骗人时,动作比嘴更致命

我们总以为 AI 撒谎靠的是话术;这篇在 3D 版《Among Us》里让 AI 当内鬼,发现它赢靠的不是嘴,是动作。研究者搭了个能看能动的沙盒,AI 内鬼要一边说话一边做动作骗过队友,结果把「说话」和「动作」拆开测,动作才是决定胜负的那一路——它会在你背后偷偷靠近、假装修船、用身体挡住你的视线。这不是你明天能用上的东西,但它第一次把 AI 的欺骗从纯文字拉进了真实世界的身体语言,给「AI 会不会在现实里骗人」这个安全议题开了条新路。

📄 原文摘要(英文)

Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新