医学伦理研究助手
前沿
论文精读

论文 · 建模 / 计算研究

RADAR:将大语言模型咨询支持锚定于ACR结构化适宜性知识库并带引用验证与拒答

J Imaging Inform Med · 2026年9月23日 · Vantaku、Avakian

问这篇
一分钟了解要点原型系统用检索增强和引用验证,减少大语言模型在影像适宜性咨询中的幻觉。结果在36个留出场景中,完整流程在33个场景中正确选择变体和首选检查(91.7%),且无一推荐错误变体;对22个临床邻近探针拒答17个,全部721条引用均可追溯到检索片段。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

Large language models (LLMs) answer medical questions fluently but fabricate statements and citations too often to be trusted unaided at the point of care. We describe RADAR (Radiology Appropriateness Decision and Advisory Resource), a prototype retrieval-augmented consult-support pipeline that grounds an LLM in a knowledge base structured like the American College of Radiology (ACR) Appropriateness Criteria at the level of individual clinical variants. The knowledge base is synthetic (twelve variants across five topics) because the official criteria are copyrighted. The design is deliberately conservative: a BM25 retriever, an absolute-score abstention gate, a structured-output generator bound by an in-prompt grounding contract, and a deterministic validator that strips any recommendation whose citation does not resolve to a retrieved chunk. We report a development set (eight scenarios, seven off-domain probes) used during parameter selection and a held-out set written after all parameters were frozen (36 scenarios, three per variant, and 24 clinically adjacent probes), evaluated in a deterministic retrieval-only mode and in the full pipeline over five independent runs at temperature 0. Retrieval alone ranked the correct variant first in 25 of 36 held-out scenarios (69.4%; 95% CI 51.9-83.7) but always placed it within the top five, and the score gate stopped only 3 of 22 clinically adjacent probes. The full pipeline (claude-sonnet-4-6, temperature 0, five runs) selected the correct variant and top procedure in every run for 33 of 36 scenarios (91.7%; 95% CI 77.5-98.2) and in 33 or 34 of 36 scenarios in each individual run; every failed run was a false abstention on a thunderclap-headache presentation, and no run of any scenario recommended a wrong variant. It refused 17 of 22 clinically adjacent probes (77.3%; 54.6-92.2). All 721 emitted citations resolved to a retrieved chunk. The findings characterize a proof of concept: lexical retrieval delivers topic routing and citation traceability, while variant selection and realistic abstention depend on the generator acting as a semantic reranker, and the residual failures are refusals rather than wrong answers. Clinical readiness will require licensed ACR content, a larger reference-standard set, and prospective evaluation.

这篇对您:
讲解或动画有问题: