论文 · 队列研究
评估基于大语言模型的自动评分在医学生语音虚拟标准化病人平台中的应用:横断面一致性研究
Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: Cross-Sectional Agreement Study
作者:Xiaoxing Gao, Xiaoming Huang, Rongrong Hu, Li Zhang, Huiting Liu, Bingqing Zhang, Chong Wei, Wei Qiu, Mengyu Zhang, Xuefeng Sun
JMIR Med Educ · 2026年9月24日 · Gao 等 10 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
BACKGROUND: Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear.
OBJECTIVE: This study aimed to assess agreement between LLM-generated and faculty ratings of history-taking and communication performance and to examine the influence of rater and case heterogeneity.
METHODS: In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, and cough). Ten blinded faculty raters scored performance (0-100 points total; 0-50 points per domain). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC [2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC).
RESULTS: Median total scores were similar for AI and faculty (median 93.0, IQR 89.0-95.0 vs median 94.0, IQR 91.0-95.0). Rater variability accounted for 37% of residual variance in faculty total scores (VPC=0.37). AI total scores were positively associated with faculty total scores (β=0.37, 95% CI 0.26-0.48; P<.001; Spearman ρ=0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1]=0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a small, not statistically significant mean bias (1.26 points, 95% CI -0.48 to 3.01; P=.16) and 95% limits of agreement from -4.95 to 7.48 (width=12.43 points), with proportional bias (β_proportional bias=-0.55; P<.001). Agreement was stronger for information gathering (β=0.46; ρ=0.49; ICC=0.54; VPC=0.23) than for communication (β=0.27; ρ=0.28; ICC=0.29; VPC=0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC=0.38).
CONCLUSIONS: LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, this method is suitable for formative use and enhanced sampling in programmatic assessment but not for independent, high-stakes summative decisions.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。