医学伦理研究助手
前沿
论文精读

论文 · 体外 / 类器官研究

多模态大语言模型用于儿童皮疹诊断的专家比较评估:临床实用性、安全性、信息质量与可读性

Children (Basel) · 2026年9月18日 · Lahut 等 3 位作者

问这篇
一分钟了解要点三种多模态大语言模型对儿童皮疹的回复多可接受,但单次作答诊断错误不少。结果约82%的临床实用性与安全性评分处于较高档,但15个病例中至少有一个平台漏诊参考诊断的有9个,三个平台都漏诊卟啉症。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要)…

已等待 0 秒大约需要 10–20 秒

可以先看别的,做好了会自动出现在这里。

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

Background/Objectives: Multimodal large language models (LLMs) can interpret clinical text and images, but their performance in pediatric rash assessment remains uncertain. This study compared the clinical utility, safety, information quality, diagnostic correctness, and readability of ChatGPT, Gemini and Grok. Methods: Fifteen content-validated pediatric rash vignettes with brief histories and anonymized photographs were submitted once to each platform using a standardized zero-shot prompt. Three pediatricians blinded to platform identity independently rated the 45 responses using a five-point Clinical Utility and Safety (CUS) scale and a five-item modified DISCERN instrument. Diagnostic correctness was assessed descriptively; platform comparisons used Friedman tests with Bonferroni-adjusted Wilcoxon tests when appropriate. Results: Overall, 82.2% of CUS ratings were in categories 4-5 and 83.0% of modified DISCERN scores were ≥20/25; no rating was assigned to CUS category 1. Gemini and Grok had descriptively higher expert ratings than ChatGPT, but CUS did not differ significantly across platforms (p = 0.157), and although modified DISCERN differed globally (p = 0.038), no pairwise comparison remained significant after adjustment. In the single-query diagnostic assessment, at least one platform missed the reference diagnosis in 9/15 vignettes, and all three missed porphyria. Gemini generated the longest responses, whereas Grok produced the most linguistically complex text; neither response length nor readability was associated with expert ratings. Conclusions: The three multimodal LLMs produced predominantly clinically acceptable responses, but performance varied by vignette and platform. Because each vignette-platform combination was sampled once, diagnostic findings represent single-response observations rather than stable platform accuracy estimates. Clinical verification remains necessary.

这篇对您:
讲解或动画有问题: