医学伦理研究助手
前沿
论文精读

论文 · 队列研究

评估基于大语言模型的自动评分在医学生语音虚拟标准化病人平台中的应用:横断面一致性研究

JMIR Med Educ · 2026年9月24日 · Gao 等 10 位作者

问这篇
一分钟了解要点比较AI与教师对医学生问诊沟通的评分,一致性中等。结果AI与教师的总分中位数相近(93.0对94.0),但教师评分本身存在较大差异,占残差方差的37%。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

BACKGROUND: Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear.

摘要第 2 段问这一段

OBJECTIVE: This study aimed to assess agreement between LLM-generated and faculty ratings of history-taking and communication performance and to examine the influence of rater and case heterogeneity.

摘要第 3 段问这一段

METHODS: In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, and cough). Ten blinded faculty raters scored performance (0-100 points total; 0-50 points per domain). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC [2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC).

摘要第 4 段问这一段

RESULTS: Median total scores were similar for AI and faculty (median 93.0, IQR 89.0-95.0 vs median 94.0, IQR 91.0-95.0). Rater variability accounted for 37% of residual variance in faculty total scores (VPC=0.37). AI total scores were positively associated with faculty total scores (β=0.37, 95% CI 0.26-0.48; P<.001; Spearman ρ=0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1]=0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a small, not statistically significant mean bias (1.26 points, 95% CI -0.48 to 3.01; P=.16) and 95% limits of agreement from -4.95 to 7.48 (width=12.43 points), with proportional bias (β_proportional bias=-0.55; P<.001). Agreement was stronger for information gathering (β=0.46; ρ=0.49; ICC=0.54; VPC=0.23) than for communication (β=0.27; ρ=0.28; ICC=0.29; VPC=0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC=0.38).

摘要第 5 段问这一段

CONCLUSIONS: LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, this method is suitable for formative use and enhanced sampling in programmatic assessment but not for independent, high-stakes summative decisions.

这篇对您:
讲解或动画有问题: