AI裁判被图片干扰,却分不清图片内容
我们越来越习惯让AI当裁判,替人类打分、标注数据。这篇论文发现一个反直觉的事:给AI裁判配一张无关图片,它的判断就会乱——但乱的方向跟图片内容没关系。研究者设计了200个句子,每个句子配三种图:符合句意的、误导的、或者没有图。规则明确说只看句子,但13个AI裁判里,配了图之后平均有约20%的答案被改变,而删掉“忽略图片”这条指令才改变11.6%。更怪的是,两张内容相反的图造成的改变幅度几乎一样,而且只有37%的改动是朝着图片暗示的方向去的——也就是说,AI被“有图”这件事干扰了,而不是被“图里是什么”干扰。这提醒我们:AI裁判的评分里,藏着评测环境的影子,你测出来的可能不是模型能力,而是配置的偶然。它不是你明天能用上的,但如果你在依赖AI做内容审核或数据标注,值得知道这个坑。
📄 原文摘要(英文)
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.