医学伦理研究助手
前沿
论文精读

论文

可读性与质量的悖论:比较ChatGPT、Gemini和Perplexity对儿童胸痛查询的回答

PLoS One · 2026年9月25日 · Hikmet Kıztanır、Ebru Çetin Özbek

问这篇
一分钟了解要点三种AI聊天机器人回答儿童胸痛问题,可读性差但Perplexity质量评分最高。结果用六种可读性公式评估发现,三种模型的可读性均显著高于六年级阅读水平门槛,即偏难;其中ChatGPT和Gemini相对易读,Perplexity最学术化。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

Pediatric chest pain represents a frequently encountered symptom in emergency departments and outpatient clinics throughout childhood, often generating substantial parental anxiety. This study aims to comparatively analyze the readability, information quality, and scientific reliability of textual contents generated by the artificial intelligence (AI) chatbots ChatGPT, Gemini, and Perplexity regarding pediatric chest pain, utilizing multidimensional analytical indices. Out of the top 25 queries with the highest global search volume on Google Trends, 17 unique keywords meeting the predefined inclusion criteria were filtered on June 1, 2026. These inquiries were directed to all three AI platforms in distinct, independent user sessions, yielding a total of 51 textual responses. Linguistic readability levels were calculated across six separate digital interfaces using the FKGL, FRES and other formulas, and the results were benchmarked against the sixth-grade reading level" threshold. Scientific reliability was audited via the Modified DISCERN and JAMA benchmarks, while content quality was examined using the GQS and EQIP instruments. The median readability scores computed across all three AI models were found to be statistically and significantly above the targeted sixth-grade comprehension threshold, indicating a high level of difficulty (p < 0.001). Post-hoc pairwise evaluations revealed that ChatGPT and Gemini offered linguistically more accessible and structurally less complex textual architectures; conversely, Perplexity exhibited a significantly more difficult and academic linguistic framework (p < 0.0167). On the other hand, regarding content quality and source credibility, Perplexity demonstrated remarkably superior and more optimized scores across all evaluated instruments, namely GQS (p = 0.002, p = 0.001), JAMA (p < 0.001, p < 0.001), mDISCERN (p = 0.005, p = 0.001), and EQIP (p = 0.003, p < 0.001), compared directly to both ChatGPT and Gemini, respectively. Although generative AI formulations harbor considerable potential to provide extensive data on pediatric chest pain, the sophisticated, university-level linguistic architecture of these texts poses a substantial access barrier for individuals who lack proficient digital health literacy skills. While Perplexity achieved the performance closest to the "gold standard" in terms of informational accuracy, no model is currently mature enough to substitute for a professional medical consultation due to observed citation biases and omissions. It is of paramount importance that clinicians actively guide families to scrutinize online health data through a critical lens.

这篇对您:
讲解或动画有问题: