AI Pulse
📄 论文解读

大脑扫描能读出你听到的话,但靠的不是声音本身

我们以为 AI 从脑电波里还原你听到的话,靠的是声音本身;这篇发现,它真正抓的是「叙事结构」——连贯的故事比随机单词更容易被还原。研究者用脑磁图(MEG)记录听故事时的大脑活动,训练 AI 把脑信号匹配到对应的音频片段,在 1005 个候选中选对的准确率达 39.75%,而解码器参数少了 20 倍。他们还给模型装上「可解释的镜头」:把脑信号映射回大脑皮层,发现左脑负责高频节律,右脑没有;再逐个遮挡刺激特征,发现沉默、音强、元音和声音起始贡献最大。最反直觉的是:把随机单词列表的脑信号换成叙事时的脑信号,还原率反而更高——说明连贯的叙事让大脑活动携带更多可恢复的信息。这不是你明天能用的技术,但它告诉我们:大脑对「故事」的编码,比我们以为的更特殊。

📄 原文摘要(英文)

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts. On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss. Together, source mapping and input interventions reveal what drives retrieval.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新