论文
评估大语言模型用于心肌梗死公共卫生教育:信息质量、透明度与可读性的比较研究
Evaluating large language models for myocardial infarction public health education: a comparative study on information quality, transparency and readability
作者:Tailong Lv, Wenkai Bao, Shudi Li, Cong Sun, Shouqiang Chen, Menghe Zhang
Front Public Health · 2026年9月10日 · Lv 等 6 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
OBJECTIVE: Myocardial infarction (MI) is an acute, life-threatening cardiovascular disease, and high-quality, accessible public health education is vital for emergency management. This study systematically evaluates the quality, transparency, clinical accuracy, patient safety, and readability of information generated by different large language models (LLMs) in responding to MI-related public inquiries.
METHODS: Twenty-five representative MI patient education questions were submitted to Gemini 3.5 Flash, Claude Opus 4.8, and ChatGPT 5.5. The generated information was independently evaluated by two cardiologists using four validated tools (DISCERN, EQIP, GQS, and JAMA) alongside a strict clinical safety assessment. Text readability was concurrently assessed utilizing six established metrics (FRES, ARI, GFI, CLI, FKGL, and SMOG).
RESULTS: Significant variations were observed in the quality, transparency and readability of information generated by the evaluated LLMs. Regarding quality and transparency, significant overall differences were noted among models in DISCERN (p < 0.001) and EQIP (p < 0.001) scores, whereas no significant differences were found in GQS and JAMA benchmarks. For DISCERN, Claude achieved significantly higher scores than both Gemini and ChatGPT. In the EQIP assessment of completeness and clarity, Claude and Gemini scored significantly higher than ChatGPT, though all models attained a "good" rating. Additionally, all models exhibited poor performance on the JAMA benchmark, indicating critical deficits in information transparency. Notably, LLMs sometimes generated incomplete, incorrect, or even potentially harmful information. Regarding readability, although Claude generated relatively more comprehensible text, all models failed to meet the recommended sixth-grade reading benchmark, indicating high reading difficulty.
CONCLUSION: While LLMs can generate structurally clear and logically coherent foundational content for MI-related queries, they occasionally produce clinically inappropriate directives. Furthermore, the texts generated by these models are overly complex, creating substantial reading barriers for the general public. Consequently, under zero-shot and English-language testing conditions, the current LLMs are not yet capable as standalone health education tools for MI.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。