论文
可读性与质量的悖论:比较ChatGPT、Gemini和Perplexity对儿童胸痛查询的回答
The readability and quality paradox: Comparing ChatGPT, Gemini, and Perplexity outputs on pediatric chest pain queries
作者:Hikmet Kıztanır, Ebru Çetin Özbek
PLoS One · 2026年9月25日 · Hikmet Kıztanır、Ebru Çetin Özbek
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
Pediatric chest pain represents a frequently encountered symptom in emergency departments and outpatient clinics throughout childhood, often generating substantial parental anxiety. This study aims to comparatively analyze the readability, information quality, and scientific reliability of textual contents generated by the artificial intelligence (AI) chatbots ChatGPT, Gemini, and Perplexity regarding pediatric chest pain, utilizing multidimensional analytical indices. Out of the top 25 queries with the highest global search volume on Google Trends, 17 unique keywords meeting the predefined inclusion criteria were filtered on June 1, 2026. These inquiries were directed to all three AI platforms in distinct, independent user sessions, yielding a total of 51 textual responses. Linguistic readability levels were calculated across six separate digital interfaces using the FKGL, FRES and other formulas, and the results were benchmarked against the sixth-grade reading level" threshold. Scientific reliability was audited via the Modified DISCERN and JAMA benchmarks, while content quality was examined using the GQS and EQIP instruments. The median readability scores computed across all three AI models were found to be statistically and significantly above the targeted sixth-grade comprehension threshold, indicating a high level of difficulty (p < 0.001). Post-hoc pairwise evaluations revealed that ChatGPT and Gemini offered linguistically more accessible and structurally less complex textual architectures; conversely, Perplexity exhibited a significantly more difficult and academic linguistic framework (p < 0.0167). On the other hand, regarding content quality and source credibility, Perplexity demonstrated remarkably superior and more optimized scores across all evaluated instruments, namely GQS (p = 0.002, p = 0.001), JAMA (p < 0.001, p < 0.001), mDISCERN (p = 0.005, p = 0.001), and EQIP (p = 0.003, p < 0.001), compared directly to both ChatGPT and Gemini, respectively. Although generative AI formulations harbor considerable potential to provide extensive data on pediatric chest pain, the sophisticated, university-level linguistic architecture of these texts poses a substantial access barrier for individuals who lack proficient digital health literacy skills. While Perplexity achieved the performance closest to the "gold standard" in terms of informational accuracy, no model is currently mature enough to substitute for a professional medical consultation due to observed citation biases and omissions. It is of paramount importance that clinicians actively guide families to scrutinize online health data through a critical lens.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。