医学伦理研究助手
前沿
论文精读

论文 · 体外 / 类器官研究

开发评估大语言模型安全性与可靠性的框架:一项概念验证评估

J Med Syst · 2026年9月23日 · Liu 等 7 位作者

问这篇
一分钟了解要点用灾难性错误率和回答可重复性评估6个大语言模型,发现总体准确率掩盖了安全性差异。结果总体得分掩盖了巨大差异,锚定偏倚失败率相差近4倍,罕见病识别率相差6倍,灾难性错误率呈两极分化;评分中等的GPT-5.1可重复性最差且落入高灾难性错误组。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

Large language models (LLMs) are entering clinical decision support faster than methodology can characterise their safety. Aggregate accuracy treats all errors as interchangeable and cannot support safe deployment under Software as a Medical Device (SaMD) and EU AI Act frameworks. To develop and demonstrate a framework for evaluating large language model safety and reliability using catastrophic-failure frequency and response reproducibility, incorporating a pre-specified error taxonomy, difficulty-stratified analysis, and two-layer response consistency. The framework was applied to 54 diagnostically challenging emergency cases from non-public institutional records. Six LLMs were accessed via application programming interfaces (APIs) (text-only, zero-shot, defaults; August-December 2025) and queried three times each, producing 972 physician-scored responses. Aggregate scores concealed heterogeneity: anchoring-bias failure (score ≤ 3) varied nearly four-fold (6.9%-26.4%) and rare-disease recognition six-fold (6.1%-37.9%); catastrophic (dangerous-recommendation, score ≤ 2) rates were two-tiered (1.5% for the safest two vs. 6.0% for the rest; p < 0.001), though within-tier ranks were inseparable (2-12 per model). Despite a middle-tier mean, GPT-5.1 was least reproducible (within-case SD 2.69) and fell within the higher catastrophic-failure tier; across models, 31 of 34 dangerous combinations were stochastic, not systematic-a failure mode hidden by aggregate metrics. Open-source DeepSeek R1 was statistically equivalent to Gemini 3 Pro within a 1.5-point margin by two one-sided tests (TOST: Δ - 0.09; 90% CI - 0.47 to 0.30; p < 0.001). In this single-centre proof-of-concept, aggregate accuracy was insufficient for safety characterisation: models with indistinguishable mean accuracy carried different catastrophic-failure tiers and reproducibility profiles. Error-type and response-consistency profiling may inform model selection, ensemble design, and conformity assessment.

这篇对您:
讲解或动画有问题: