AI会诊癌症,但结论还差得远
让 AI 像真实医院的多学科会诊那样讨论癌症病例——10 个专科角色、6 万多分钟真实会诊录音转写成的 611 个病例,是目前最接近临床现实的测试。结果:最强的模型在回答专科问题上只拿到 3.43/5 的临床等价分,在跟真实会诊结论对齐上只有 2.78/5。也就是说,AI 能像模像样地讨论,但离真正替医生拍板还差一大截。它不是你明天能用上的工具,但这份基准让「AI 参与癌症决策」从口号变成了可量化的进度条。
📄 原文摘要(英文)
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.