让AI读胸片:把整份报告拆成一句句,零样本也能看懂
现在的医学影像AI大多靠“看图配文”训练,但报告又长又密,模型很难直接拿来回答“这张胸片有没有肺结节”。这篇把报告拆成一句句临床描述,再让大模型把意思相近的句子归并成抽象概念,既扩充了训练样本,又避免把“不同病人说同一句话”误判成反例。结果是在没见过的任务上,零样本表现超过此前所有多任务方法。它不是你明天就能用的工具,但指向一个趋势:医学AI正在从“背答案”转向“真读片”。
📄 原文摘要(英文)
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.