利用大语言模型进行根本原因分析以加强肿瘤放射治疗中的患者安全监测:一项概念验证研究
Augmenting patient safety surveillance in radiation oncology with large language model-based root cause analysis: A proof-of-concept study
放射治疗中的事故需要做根本原因分析(找出事件发生的深层原因),但这项工作费时费力。研究者把放射肿瘤学事故学习系统中19份公开报告的背景与事件概述部分,输入四个先进的大语言模型(Gemini 2.5 Pro、GPT-4o、o3和Grok 3),让它们按美国医学物理学家协会的指南生成根本原因、经验教训和改进建议。五位有资质的医学物理师参与评审。结果显示模型整体表现合格,但都会出现不同程度的幻觉(编造不存在的内容),比例从11%到61%不等;Gemini 2.5 Pro综合得分最高。作者认为大语言模型可作为辅助工具支持事故分析,但幻觉问题需要警惕。
为什么推荐给您:把大语言模型用于放射治疗事故根本原因分析属较新应用场景,但仅为概念验证、样本仅19例。
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
We investigated the potential utility of large language models (LLMs) in supporting patient safety efforts. Specifically, we evaluated the reasoning capabilities of LLMs in performing root cause analysis (RCA) of radiation oncology incidents using narrative reports from the Radiation Oncology Incident Learning System (RO-ILS). We prompted four state-of-the-art LLMs, Gemini 2.5 Pro, GPT-4o, o3, and Grok 3, with the "Background and Incident Overview" sections from 19 publicly available RO-ILS cases. Each model was instructed to perform RCA and generate root causes, lessons learned, and suggested actions using a standardized prompt based on AAPM RCA guidelines. Model outputs were evaluated using a combination of objective semantic similarity metrics (cosine similarity via Sentence Transformer), semi-subjective assessments (precision, recall, F1-score, expert-adjudicated PPV (Positive Predictive Value), hallucination rate and performance criteria including relevance, comprehensiveness, quality of justification and quality of solution), and subjective ratings (reasoning quality and overall performance) by five board-certified medical physicists. LLMs demonstrated satisfactory performance across evaluation metrics. All models exhibited some degree of hallucination, ranging from 11% to 61%. All the evaluated LLMs demonstrated comparable baseline capabilities in objective causal extraction, and Gemini 2.5 Pro exhibited the highest overall performance score among 4 models. Statistically significant differences were observed among models in expert-adjudicated PPV, hallucination rate, and subjective ratings (p < 0.05). LLMs delivered promising results as assistive tools for RCA in radiation oncology, with the ability to generate relevant and accurate analyses aligned with expert expectations. LLMs may support incident analysis and contribute to quality improvement efforts to advance patient safety in clinical radiation oncology practice.