论文
大语言模型回答肝硬化营养公共问题的表现:一项比较研究
Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study
作者:Junzhen Li, Yingjie Wu, Man Yang, Lihua Ren, Jian Qi, Hui Ouyang
Nutr Hosp · 2026年9月16日 · Li 等 6 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
BACKGROUND: large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality, and readability.
METHODS: in this cross-sectional comparative study, 50 public facing questions about nutrition in cirrhosis were submitted once to each of four LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, and Gemini 3.1 Pro Thinking) between May 10 and May 16, 2026. Three raters independently assessed 200 responses for safety, accuracy, empathy, and information quality. Readability was assessed with six indices.
RESULTS: a total of 200 responses were evaluated. Safety coding showed high inter-rater agreement (Fleiss' kappa = 0.864; 95 % CI, 0.755-0.958). ICC values for other manually scored outcomes ranged from 0.860 to 0.892. Potentially unsafe responses occurred in all models, ranging from 6.0 % to 12.0 %, with no significant difference across models (raw p = 0.599). Accuracy differed significantly across models (raw p = 0.003), with the highest median score for ChatGPT 5.5. Empathy also differed significantly (raw p < 0.001), with higher median scores for Claude Opus 4.7 and Gemini 3.1 Pro Thinking. DISCERN, EQIP, JAMA and all readability indices differed significantly across models, while GQS did not.
CONCLUSION: current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。