论文
评估大型语言模型(ChatGPT、DeepSeek、Gemini)回答患者关于脂肪水肿问题的准确性与可重复性
Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions
作者:Rabia Sanır, Esra Nur Türkmen, Esra Giray, Figen Ayhan, Gülseren Akyüz
Phlebology · 2026年9月23日 · Rabia Sanır 等 5 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered information require thorough evaluation, especially in rare diseases like lipedema. Therefore, this study aimed to evaluate the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to questions frequently asked by patients with lipedema.MethodsThis cross-sectional study assessed the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to 25 commonly asked lipedema-related questions. Each model was queried twice in separate sessions, and answers were evaluated by three independent experts using a four-point rating scale. To ensure the objectivity and consistency of expert evaluations, inter-rater agreement was assessed using Cohen's kappa coefficient.ResultsDeepSeek achieved the highest proportion of comprehensive and correct responses (72%), followed by Gemini (64%) and ChatGPT (56%). Accuracy varied across content categories, with notable limitations particularly in treatment, follow-up, and maintenance questions. Reproducibility analysis revealed that DeepSeek produced the most consistent responses across sessions, while ChatGPT and Gemini showed more variability, particularly in treatment and quality-of-life questions. Cohen's kappa values indicated high inter-rater agreement overall, with perfect agreement in some categories for ChatGPT and DeepSeek.ConclusionsLLMs can provide generally accurate and consistent responses to patient-centered questions about lipedema, particularly in areas related to general information and diagnosis. However, reduced accuracy and reproducibility in complex clinical domains suggest that expert oversight is essential when using these tools for patient education.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。