论文
五种大型语言模型驱动的聊天机器人在44个专家审定家长提问(儿童维生素D缺乏)上的横断面比较评估
Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency
作者:Ping Shi, Tian Zhou, Qiao Nie, Baicheng Tao, Xiaolu Li, Mei Zhu
Front Pediatr · 2026年9月9日 · Shi 等 6 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
BACKGROUND: Parents and caregivers increasingly use generative artificial intelligence chatbots for child health information. Pediatric vitamin D deficiency is a clinically relevant topic because advice about supplementation, testing, rickets, high-risk children, toxicity, and emergency symptoms can influence caregiver decisions.
OBJECTIVE: To compare the safety, medical accuracy, empathy, information reliability, educational quality, transparency, global quality, and readability of five large language model-driven chatbots when answering questions about pediatric vitamin D deficiency.
METHODS: This cross-sectional comparative study was reported with reference to CHART. An expert-curated 44-question set was selected from a 92-question candidate pool informed by search trends, caregiver-facing sources, clinical guidelines, and expert discussion. Each of the 44 questions was submitted once to each of five chatbot services-ChatGPT-5.5, Gemini 3.1 Pro, Qianwen 3.6-Plus, DeepSeek V4, and Doubao-Seed-2.0 Pro-using the same parent-oriented instruction. Responses were assessed for safety, accuracy, empathy, DISCERN, EQIP, JAMA criteria, GQS, and readability. Paired question-level differences were analyzed using Friedman tests, Kendall's W, and Cochran's Q; results were interpreted as a single-run, time-specific snapshot.
RESULTS: Inter-rater agreement was good to excellent (Fleiss' kappa = 0.842 for safety; ICCs 0.846-0.914 for other subjective metrics). Twenty of 220 responses (9.1%) were unsafe; safety did not differ significantly across models (Cochran's Q = 3.704, df = 4, P = 0.448). Significant inter-model differences were observed in accuracy, empathy, DISCERN, EQIP, JAMA, GQS, and all readability indices. ChatGPT had the highest median accuracy [5.00 (4.00, 5.00)] and DISCERN score [69.60 (65.25, 72.40)]; DeepSeek and Doubao had the highest empathy scores [5.00 (4.80, 5.00)]. ChatGPT and Doubao shared the highest EQIP median (86.00), Doubao had the highest GQS [5.00 (4.00, 5.00)], and Gemini had the highest FRES. JAMA scores were low across models.
CONCLUSIONS: In this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。