AI Pulse
📄 论文解读

人听不见的低频声,能让AI语音助手直接失灵

你听不见的低频噪音,正在悄悄让AI语音助手“耳聋”。研究者发现,大音频语言模型(LALM)能接收人耳听不到的低频信号,并受其干扰——他们用一段通用波形模板,在6个主流模型上把任务准确率最高拉低67个百分点,而人类几乎察觉不到异常(可听度评分1.33,接近纯净音频的1.17)。更关键的是,他们同时给出了一道防线:检测到低频分布偏移时,让模型请求重新录音,能把被攻击后的平均准确率从28.5%拉回46.1%。这不是你明天能用的功能,但它揭示了一个真实存在的安全盲区:当AI开始“听”得比人更广,攻击面也随之扩大。

📄 原文摘要(英文)

Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新