医学伦理研究助手
前沿
论文精读

论文 · 建模 / 计算研究

大语言模型诊断进食障碍中的隐性偏见:实验性情景研究

JMIR AI · 2026年9月23日 · McCalla 等 7 位作者

问这篇
一分钟了解要点相同临床表现下,LLM因种族等人口学特征给出系统性不同的精神科诊断。结果对照情景几乎一致诊断为AN(均值100%),模糊情景差异大(均值23.6%)。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

BACKGROUND: Large language models (LLMs) are increasingly deployed in mental health applications, yet growing evidence suggests they encode algorithmic biases that influence clinical outputs. Because these models now mediate patient-facing decisions, such biases carry the potential for direct harm. Whether they systematically affect psychiatric diagnosis across demographic groups remains underexplored.

摘要第 2 段问这一段

OBJECTIVE: This study aims to examine whether LLMs exhibit implicit demographic biases when generating psychiatric diagnoses.

摘要第 3 段问这一段

METHODS: We developed 1152 synthetic clinical vignettes using a matched-pair design that manipulated gender, race and ethnicity, age, socioeconomic status, English proficiency, and urbanicity while holding clinical content constant. Vignettes were divided into control (unambiguous anorexia nervosa [AN]) and ambiguous conditions designed to permit differential diagnosis. Ten LLM configurations across 5 model families were tested.

摘要第 4 段问这一段

RESULTS: Control vignettes produced near-unanimous AN diagnoses (mean 100%, SD 0.1%), while ambiguous vignettes elicited greater variability (mean 23.6%, SD 10.1%). Intermodel agreement was moderate for ambiguous vignettes (Fleiss κ=0.410, 95% CI 0.397-0.422). Mixed-effects logistic regression with LLM as a random intercept revealed significant demographic biases: Black patients were over 6 times more likely to receive a major depressive disorder (MDD) diagnosis than White patients with identical presentations (odds ratio [OR] 6.09, 95% CI 5.13-7.24), Latine patients were over 9 times more likely (OR 9.57, 95% CI 8.00-11.45), and Asian patients were nearly 3 times more likely to receive an AN diagnosis (OR 2.88, 95% CI 2.44-3.42). Female patients were less likely than males to be diagnosed with AN (OR 0.43, 95% CI 0.37-0.49).

摘要第 5 段问这一段

CONCLUSIONS: These findings demonstrate that LLMs exhibit systematic demographic biases in psychiatric diagnosis even when clinical content is held constant, revealing measurable patterns that can inform improvements to training data, model architecture, and clinical deployment frameworks.

这篇对您:
讲解或动画有问题: