医学伦理研究助手
前沿
论文精读

论文 · 队列研究

大语言模型遇上妇科超声:提升附件肿块的鉴别能力

J Imaging · 2026年9月21日 · Soccio 等 9 位作者

问这篇
一分钟了解要点比较聊天机器人与专家判读附件肿块良恶性的准确性。结果以病理检查为金标准,专家判断准确率最高,为 87.3%;聊天机器人准确率为 73.3% 和 74.7%,敏感性约 72%–74%,特异性约 75%–76%。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

Ovarian cancer (OC) is the second most common gynecological malignancy and remains one of the leading causes of gynecological cancer-related mortality worldwide. A major clinical challenge is the lack of an accurate and widely applicable strategy for identifying patients at high risk of malignancy at an early stage. In this context, artificial intelligence (AI) has emerged as a promising tool to improve diagnostic performance. Among AI technologies, large language models (LLMs) have recently shown considerable potential in healthcare applications. In this study, we evaluated the diagnostic performance of ChatGPT (GPT-5) in classifying 300 adnexal masses as benign or malignant and compared its performance with that of the IOTA Simple Rules, the ADNEX model, and expert subjective assessment. We also assessed ChatGPT's ability to predict the most likely histological diagnosis for each lesion. All adnexal masses were described using the International Ovarian Tumor Analysis (IOTA) terminology, and histopathological examination served as the reference standard. Our findings showed that expert subjective assessment achieved the highest overall diagnostic performance for both benign/malignant classification (accuracy 87.3%; 95% CI, 83.0-90.9%) and prediction of the presumed histological diagnosis. ChatGPT A and ChatGPT B reached a sensitivity of 72.3% and 73.5%, a specificity of 74.5% and 75.9%, a positive predictive value of 75.2% and 76.5%, and a negative predictive value of 71.5% and 72.8%, respectively (inconclusive responses counted as misclassifications), with an overall accuracy of 73.3% and 74.7%. After adequate validation, large language models might complement existing decision-support tools for less experienced examiners, without replacing expert evaluation. Their ease of use and reliance on standardized ultrasound descriptors make them accessible to ultrasonographers with varying levels of expertise.

这篇对您:
讲解或动画有问题: