医学伦理研究助手

利用大语言模型进行根本原因分析以加强肿瘤放射治疗中的患者安全监测:一项概念验证研究

Augmenting patient safety surveillance in radiation oncology with large language model-based root cause analysis: A proof-of-concept study

PLOS Digit Health · 2026 年 9 月 25 日 · Yuntao Wang, Mariluz De Ornelas, Matthew T Studenski 等 6 人

体外 / 类器官研究
在聊天里讨论
一分钟了解
用四种大语言模型对19例放射治疗事故报告做根本原因分析,结果与专家判断基本一致但存在幻觉。

放射治疗中的事故需要做根本原因分析(找出事件发生的深层原因),但这项工作费时费力。研究者把放射肿瘤学事故学习系统中19份公开报告的背景与事件概述部分,输入四个先进的大语言模型(Gemini 2.5 Pro、GPT-4o、o3和Grok 3),让它们按美国医学物理学家协会的指南生成根本原因、经验教训和改进建议。五位有资质的医学物理师参与评审。结果显示模型整体表现合格,但都会出现不同程度的幻觉(编造不存在的内容),比例从11%到61%不等;Gemini 2.5 Pro综合得分最高。作者认为大语言模型可作为辅助工具支持事故分析,但幻觉问题需要警惕。

为什么推荐给您:把大语言模型用于放射治疗事故根本原因分析属较新应用场景,但仅为概念验证、样本仅19例。

讲解深度:

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

摘要Abstract

摘要第 1 段问这一段

We investigated the potential utility of large language models (LLMs) in supporting patient safety efforts. Specifically, we evaluated the reasoning capabilities of LLMs in performing root cause analysis (RCA) of radiation oncology incidents using narrative reports from the Radiation Oncology Incident Learning System (RO-ILS). We prompted four state-of-the-art LLMs, Gemini 2.5 Pro, GPT-4o, o3, and Grok 3, with the "Background and Incident Overview" sections from 19 publicly available RO-ILS cases. Each model was instructed to perform RCA and generate root causes, lessons learned, and suggested actions using a standardized prompt based on AAPM RCA guidelines. Model outputs were evaluated using a combination of objective semantic similarity metrics (cosine similarity via Sentence Transformer), semi-subjective assessments (precision, recall, F1-score, expert-adjudicated PPV (Positive Predictive Value), hallucination rate and performance criteria including relevance, comprehensiveness, quality of justification and quality of solution), and subjective ratings (reasoning quality and overall performance) by five board-certified medical physicists. LLMs demonstrated satisfactory performance across evaluation metrics. All models exhibited some degree of hallucination, ranging from 11% to 61%. All the evaluated LLMs demonstrated comparable baseline capabilities in objective causal extraction, and Gemini 2.5 Pro exhibited the highest overall performance score among 4 models. Statistically significant differences were observed among models in expert-adjudicated PPV, hallucination rate, and subjective ratings (p < 0.05). LLMs delivered promising results as assistive tools for RCA in radiation oncology, with the ability to generate relevant and accurate analyses aligned with expert expectations. LLMs may support incident analysis and contribute to quality improvement efforts to advance patient safety in clinical radiation oncology practice.

从这篇论文记下的摘录
在“讲解”“原文”里选中文字,会出现“记到笔记”按钮(电脑上在文字旁边,手机上在屏幕最下面);记下的内容会按笔记本整理,也会列在这里。
讲解或动画有问题?告诉我: