医学伦理研究助手
前沿
论文精读

论文 · 建模 / 计算研究

多模态大语言模型的放射学教学病例基准有多稳健?对18个模型的暴露机会、推理配置与纯文本可解性审计

J Imaging Inform Med · 2026年9月22日 · Fan 等 6 位作者

问这篇
一分钟了解要点审计放射学基准发现,模型得分受训练数据暴露与配置影响,不能当作临床诊断能力证据。结果模型准确率61.5%–80.8%,没有模型超过所有人类阅片者;队列差异不能用训练暴露解释,分辨率改动影响不超过3个百分点。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要)…

已等待 0 秒大约需要 10–20 秒

可以先看别的,做好了会自动出现在这里。

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

Multimodal large language models (MLLMs) are benchmarked on radiology multiple-choice questions. We audited how training-data exposure opportunity, image resolution, inference configuration, and text-only solvability affected interpretation of a teaching-case benchmark. Eighteen MLLMs and 2 original reference readers answered 260 questions from 51 AuntMinnie cases; 5 additional readers completed them under a unified interface. Analyses included Holm-adjusted McNemar tests, case-cluster bootstrapping, model-cutoff matching, a case-random-intercept logistic model, cross-cohort configuration re-runs, and text-only ablation. Model accuracy ranged from 61.5% to 80.8%, and the 7 readers ranged from 66.5% to 84.2%. No model exceeded all readers; 17 fell within the reader range and 1 below. Eight models were significantly less accurate than Reader 1 after reader-specific Holm adjustment (6 when all 36 comparisons were adjusted together); none differed significantly from Reader 2. A cohort/publication-period association remained after item adjustment (odds ratio, 4.67; 95% Laplace interval, 1.82-11.98), but 3 within-cohort cutoff comparisons showed no exposure-consistent advantage. Resolution changes were within 3 percentage points in both cohorts, and the configuration-by-version interaction was not significant. Among 5 post hoc selected models, mean image contribution was 6.2 percentage points in the original cohort and 4.0 in the supplementary cohort. The benchmark did not support a single human-model ranking, a training-exposure explanation for the cohort difference, or a configuration-dependent GPT version effect. Public teaching-case benchmarks should report provenance, individual-reader calibration, multiplicity control, inference settings, and image-free performance and should not be interpreted as evidence of clinical diagnostic readiness.

这篇对您:
讲解或动画有问题: