医学伦理研究助手
前沿
论文精读

论文

大语言模型回答肝硬化营养公共问题的表现:一项比较研究

Nutr Hosp · 2026年9月16日 · Li 等 6 位作者

问这篇
一分钟了解要点四款大语言模型回答肝硬化营养问题,安全性、准确性和可读性参差不齐。结果所有模型均存在潜在不安全回答(6.0%–12.0%),准确性、同理心和可读性存在显著差异。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要)…

已等待 0 秒大约需要 10–20 秒

可以先看别的,做好了会自动出现在这里。

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

BACKGROUND: large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality, and readability.

摘要第 2 段问这一段

METHODS: in this cross-sectional comparative study, 50 public facing questions about nutrition in cirrhosis were submitted once to each of four LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, and Gemini 3.1 Pro Thinking) between May 10 and May 16, 2026. Three raters independently assessed 200 responses for safety, accuracy, empathy, and information quality. Readability was assessed with six indices.

摘要第 3 段问这一段

RESULTS: a total of 200 responses were evaluated. Safety coding showed high inter-rater agreement (Fleiss' kappa = 0.864; 95 % CI, 0.755-0.958). ICC values for other manually scored outcomes ranged from 0.860 to 0.892. Potentially unsafe responses occurred in all models, ranging from 6.0 % to 12.0 %, with no significant difference across models (raw p = 0.599). Accuracy differed significantly across models (raw p = 0.003), with the highest median score for ChatGPT 5.5. Empathy also differed significantly (raw p < 0.001), with higher median scores for Claude Opus 4.7 and Gemini 3.1 Pro Thinking. DISCERN, EQIP, JAMA and all readability indices differed significantly across models, while GQS did not.

摘要第 4 段问这一段

CONCLUSION: current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.

这篇对您:
讲解或动画有问题: