AI看视频时,耳朵会骗眼睛
多模态大模型看视频时,声音和画面会互相干扰:明明画面里没有的东西,因为音频里提到了,模型就一本正经地说“看到了”。研究者发现这不是模型乱猜,而是内部有个“中继站”——问题指令在传递时把不该用的模态信息也带上了。他们用对比法把问题状态往正确模态方向“掰”一下,不用重新训练,就能让模型少犯这类错,在测试集上准确率最高提升18个百分点。它不是你明天能用上的东西,但它解释了为什么AI会“睁眼说瞎话”,以及这类幻觉不是玄学,而是可定位、可修正的机制问题。
📄 原文摘要(英文)
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: source-confused grounding hallucination, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a question-relay mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose SECRET (SourcE-Conditioned RElay sTeering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.