论文 · 建模 / 计算研究
多模态大语言模型的放射学教学病例基准有多稳健?对18个模型的暴露机会、推理配置与纯文本可解性审计
How Robust Are Radiology Teaching-Case Benchmarks for Multimodal Large Language Models? An Audit of Exposure Opportunity, Inference Configuration, and Text-Only Solvability Across Eighteen Models
作者:Wenlong Fan, Xiaoliang Chen, Yue Lin, Xintong Song, Yang Xu, Sheng Xie
J Imaging Inform Med · 2026年9月22日 · Fan 等 6 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
Multimodal large language models (MLLMs) are benchmarked on radiology multiple-choice questions. We audited how training-data exposure opportunity, image resolution, inference configuration, and text-only solvability affected interpretation of a teaching-case benchmark. Eighteen MLLMs and 2 original reference readers answered 260 questions from 51 AuntMinnie cases; 5 additional readers completed them under a unified interface. Analyses included Holm-adjusted McNemar tests, case-cluster bootstrapping, model-cutoff matching, a case-random-intercept logistic model, cross-cohort configuration re-runs, and text-only ablation. Model accuracy ranged from 61.5% to 80.8%, and the 7 readers ranged from 66.5% to 84.2%. No model exceeded all readers; 17 fell within the reader range and 1 below. Eight models were significantly less accurate than Reader 1 after reader-specific Holm adjustment (6 when all 36 comparisons were adjusted together); none differed significantly from Reader 2. A cohort/publication-period association remained after item adjustment (odds ratio, 4.67; 95% Laplace interval, 1.82-11.98), but 3 within-cohort cutoff comparisons showed no exposure-consistent advantage. Resolution changes were within 3 percentage points in both cohorts, and the configuration-by-version interaction was not significant. Among 5 post hoc selected models, mean image contribution was 6.2 percentage points in the original cohort and 4.0 in the supplementary cohort. The benchmark did not support a single human-model ranking, a training-exposure explanation for the cohort difference, or a configuration-dependent GPT version effect. Public teaching-case benchmarks should report provenance, individual-reader calibration, multiplicity control, inference settings, and image-free performance and should not be interpreted as evidence of clinical diagnostic readiness.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。