论文 · 建模 / 计算研究
超越准确率:大语言模型在皮肤科执业考试题上的任务脆弱性与推理稳定性教育基准研究
Beyond accuracy: an educational benchmarking study of task fragility and reasoning stability of large language models on dermatology board-style questions
作者:Ahmet Uğur Atılan, Nıyazı Çetın
Cutan Ocul Toxicol · 2026年9月23日 · Ahmet Uğur Atılan、Nıyazı Çetın
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
BACKGROUND: Large language models (LLMs) are increasingly used in medical education and assessment frameworks. However, their reliability under varying task demands remains unclear. Significant gaps exist in the literature regarding the reasoning stability of these models when faced with fluctuations in question structure, difficulty, and linguistic framing. In specialized fields such as dermatology, where diagnostic precision is critical, addressing these inconsistencies is essential for the reliable integration of AI into educational training environments.
METHODS: This study was designed as an educational benchmarking analysis under controlled assessment conditions. We evaluated 4 LLMs (ChatGPT 5.0, Claude 4.5 Sonnet, Gemini 2.5 Flash, DeepSeek V3) on 136 dermatology board-style multiple-choice questions, each adapted into 4 task methods: (1) original single-best-answer format; (2) correct option replaced with "None of the answers;" (3) four correct options, requiring identification of the single incorrect option; and (4) no correct option (four distractors). Items were classified as Easy, Medium, or Hard and as Positive or Negative stems. The primary outcome was Correct Response (1/0) across 2,176 model-item-method observations. A generalized linear mixed-effects model with random intercepts for Question ID estimated the effects of model, method, difficulty, and polarity, including interactions; Holm adjustment controlled for multiple comparisons.
RESULTS: Overall accuracy was 52%. Using ChatGPT as the reference, all models showed significantly lower odds of correctness (Claude OR, 0.45; Gemini OR, 0.56; DeepSeek OR, 0.54; all p < 0.001). Task method had the largest effect, with marked reductions from Method 1 to Methods 2-4 (Method 4 OR, 0.026; p < 0.001). Negative stems lowered accuracy (OR, 0.57; p < 0.001). Method × Difficulty, Method × Polarity, and Difficulty × Polarity interactions were significant, as was the three-way interaction (p = 0.006), with the steepest losses for medium-difficulty negative items in Methods 2 and 4.
CONCLUSION: Overall, LLM performance exhibited sensitivity to task design, difficulty, and linguistic framing within this educational assessment framework. These findings suggest task fragility and warrant caution in unsupervised educational applications. Practically, this highlights the value of format-aware evaluation when considering LLMs for medical examination preparation or automated question generation, helping to align expectations with their observed educational capabilities.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。