AI Pulse
📄 论文解读

AI问诊精神科:聊天流畅≠诊断准确

AI在精神科问诊上有个反直觉的发现:聊天越顺畅,诊断反而可能越不准。研究者构建了中文精神科问诊基准LingxiDiagBench,包含1.6万段模拟医患对话,覆盖12类精神疾病。测试多个大模型后发现:模型在简单任务(如区分抑郁和焦虑)上准确率可达92%,但面对共病(抑郁+焦虑)时骤降至43%,12类鉴别诊断仅28.5%。更关键的是,动态问诊(模拟真实对话)的表现普遍差于静态分析(直接给病历),说明模型问诊时收集信息策略低效,反而干扰了判断。而且,用AI评判问诊质量与诊断准确率的相关性很弱——问得头头是道,不等于能下对诊断。这不是你明天能用的工具,但它揭示了AI医疗的一个深层问题:对话流畅性可能掩盖诊断能力的不足。

📄 原文摘要(英文)

Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely and consistent mental-health assessment. Progress in AI-assisted psychiatric diagnosis is constrained by the absence of benchmarks that simultaneously provide realistic patient simulation, clinician-verified diagnostic labels, and support for dynamic multi-turn consultation. We present LingxiDiagBench, a large-scale multi-agent benchmark that evaluates LLMs on both static diagnostic inference and dynamic multi-turn psychiatric consultation in Chinese. At its core is LingxiDiag-16K, a dataset of 16,000 EMR-aligned synthetic consultation dialogues designed to reproduce real clinical demographic and diagnostic distributions across 12 ICD-10 psychiatric categories. Through extensive experiments across state-of-the-art LLMs, we establish key findings: (1) although LLMs achieve high accuracy on binary depression--anxiety classification (up to 92.3%), performance deteriorates substantially for depression--anxiety comorbidity recognition (43.0%) and 12-way differential diagnosis (28.5%); (2) dynamic consultation often underperforms static evaluation, indicating that ineffective information-gathering strategies significantly impair downstream diagnostic reasoning; (3) consultation quality assessed by LLM-as-a-Judge shows only moderate correlation with diagnostic accuracy, suggesting that well-structured questioning alone does not ensure correct diagnostic decisions. We release LingxiDiag-16K and the full evaluation framework to support reproducible research at https://github.com/Lingxi-mental-health/LingxiDiagBench.

arXiv 原文

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新