医学伦理研究助手
前沿
论文精读

论文

评估大型语言模型(ChatGPT、DeepSeek、Gemini)回答患者关于脂肪水肿问题的准确性与可重复性

Phlebology · 2026年9月23日 · Rabia Sanır 等 5 位作者

问这篇
一分钟了解要点比较三个大模型回答脂肪水肿患者问题的准确性和一致性。结果DeepSeek回答完整且正确的比例最高(72%),Gemini为64%,ChatGPT为56%;在治疗和随访类问题上三者都较弱。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要)…

已等待 0 秒大约需要 10–20 秒

可以先看别的,做好了会自动出现在这里。

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered information require thorough evaluation, especially in rare diseases like lipedema. Therefore, this study aimed to evaluate the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to questions frequently asked by patients with lipedema.MethodsThis cross-sectional study assessed the accuracy and reproducibility of responses generated by ChatGPT, DeepSeek, and Gemini to 25 commonly asked lipedema-related questions. Each model was queried twice in separate sessions, and answers were evaluated by three independent experts using a four-point rating scale. To ensure the objectivity and consistency of expert evaluations, inter-rater agreement was assessed using Cohen's kappa coefficient.ResultsDeepSeek achieved the highest proportion of comprehensive and correct responses (72%), followed by Gemini (64%) and ChatGPT (56%). Accuracy varied across content categories, with notable limitations particularly in treatment, follow-up, and maintenance questions. Reproducibility analysis revealed that DeepSeek produced the most consistent responses across sessions, while ChatGPT and Gemini showed more variability, particularly in treatment and quality-of-life questions. Cohen's kappa values indicated high inter-rater agreement overall, with perfect agreement in some categories for ChatGPT and DeepSeek.ConclusionsLLMs can provide generally accurate and consistent responses to patient-centered questions about lipedema, particularly in areas related to general information and diagnosis. However, reduced accuracy and reproducibility in complex clinical domains suggest that expert oversight is essential when using these tools for patient education.

这篇对您:
讲解或动画有问题: