医学伦理研究助手
前沿
论文精读

论文 · 建模 / 计算研究

超越准确率:大语言模型在皮肤科执业考试题上的任务脆弱性与推理稳定性教育基准研究

Cutan Ocul Toxicol · 2026年9月23日 · Ahmet Uğur Atılan、Nıyazı Çetın

问这篇
一分钟了解要点改变题目格式后LLM准确率骤降,暴露推理脆弱性。结果总体准确率52%。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要)…

已等待 0 秒大约需要 10–20 秒

可以先看别的,做好了会自动出现在这里。

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

BACKGROUND: Large language models (LLMs) are increasingly used in medical education and assessment frameworks. However, their reliability under varying task demands remains unclear. Significant gaps exist in the literature regarding the reasoning stability of these models when faced with fluctuations in question structure, difficulty, and linguistic framing. In specialized fields such as dermatology, where diagnostic precision is critical, addressing these inconsistencies is essential for the reliable integration of AI into educational training environments.

摘要第 2 段问这一段

METHODS: This study was designed as an educational benchmarking analysis under controlled assessment conditions. We evaluated 4 LLMs (ChatGPT 5.0, Claude 4.5 Sonnet, Gemini 2.5 Flash, DeepSeek V3) on 136 dermatology board-style multiple-choice questions, each adapted into 4 task methods: (1) original single-best-answer format; (2) correct option replaced with "None of the answers;" (3) four correct options, requiring identification of the single incorrect option; and (4) no correct option (four distractors). Items were classified as Easy, Medium, or Hard and as Positive or Negative stems. The primary outcome was Correct Response (1/0) across 2,176 model-item-method observations. A generalized linear mixed-effects model with random intercepts for Question ID estimated the effects of model, method, difficulty, and polarity, including interactions; Holm adjustment controlled for multiple comparisons.

摘要第 3 段问这一段

RESULTS: Overall accuracy was 52%. Using ChatGPT as the reference, all models showed significantly lower odds of correctness (Claude OR, 0.45; Gemini OR, 0.56; DeepSeek OR, 0.54; all p < 0.001). Task method had the largest effect, with marked reductions from Method 1 to Methods 2-4 (Method 4 OR, 0.026; p < 0.001). Negative stems lowered accuracy (OR, 0.57; p < 0.001). Method × Difficulty, Method × Polarity, and Difficulty × Polarity interactions were significant, as was the three-way interaction (p = 0.006), with the steepest losses for medium-difficulty negative items in Methods 2 and 4.

摘要第 4 段问这一段

CONCLUSION: Overall, LLM performance exhibited sensitivity to task design, difficulty, and linguistic framing within this educational assessment framework. These findings suggest task fragility and warrant caution in unsupervised educational applications. Practically, this highlights the value of format-aware evaluation when considering LLMs for medical examination preparation or automated question generation, helping to align expectations with their observed educational capabilities.

这篇对您:
讲解或动画有问题: