论文 · 体外 / 类器官研究
开发评估大语言模型安全性与可靠性的框架:一项概念验证评估
Development of a Framework for Evaluating Large Language Model Safety and Reliability: a Proof-of-Concept Evaluation
作者:Fangyan Liu, Zhi Liu, Xiaolu Fei, Jingyu He, Jin Xing, Jia Li, Piu Chan
J Med Syst · 2026年9月23日 · Liu 等 7 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
Large language models (LLMs) are entering clinical decision support faster than methodology can characterise their safety. Aggregate accuracy treats all errors as interchangeable and cannot support safe deployment under Software as a Medical Device (SaMD) and EU AI Act frameworks. To develop and demonstrate a framework for evaluating large language model safety and reliability using catastrophic-failure frequency and response reproducibility, incorporating a pre-specified error taxonomy, difficulty-stratified analysis, and two-layer response consistency. The framework was applied to 54 diagnostically challenging emergency cases from non-public institutional records. Six LLMs were accessed via application programming interfaces (APIs) (text-only, zero-shot, defaults; August-December 2025) and queried three times each, producing 972 physician-scored responses. Aggregate scores concealed heterogeneity: anchoring-bias failure (score ≤ 3) varied nearly four-fold (6.9%-26.4%) and rare-disease recognition six-fold (6.1%-37.9%); catastrophic (dangerous-recommendation, score ≤ 2) rates were two-tiered (1.5% for the safest two vs. 6.0% for the rest; p < 0.001), though within-tier ranks were inseparable (2-12 per model). Despite a middle-tier mean, GPT-5.1 was least reproducible (within-case SD 2.69) and fell within the higher catastrophic-failure tier; across models, 31 of 34 dangerous combinations were stochastic, not systematic-a failure mode hidden by aggregate metrics. Open-source DeepSeek R1 was statistically equivalent to Gemini 3 Pro within a 1.5-point margin by two one-sided tests (TOST: Δ - 0.09; 90% CI - 0.47 to 0.30; p < 0.001). In this single-centre proof-of-concept, aggregate accuracy was insufficient for safety characterisation: models with indistinguishable mean accuracy carried different catastrophic-failure tiers and reproducibility profiles. Error-type and response-consistency profiling may inform model selection, ensemble design, and conformity assessment.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。