大语言模型在老年精神科评估中的系统性失败模式:范围综述
Systemic Failure Modes of Large Language Models in Geriatric Psychiatric Assessment: A Scoping Review
大语言模型在医学考试中表现优异,但在真实老年精神科评估中的可靠性存疑,因为谵妄、痴呆和晚发抑郁常与衰弱、多重用药和漫长病史交织。研究系统检索2023年至2025年间的实证评估,纳入47项研究,归纳出四种反复出现的失败模式:长程推理中的诊断不稳定;对抗性输入下的脆弱性,包括幻觉和错误推理;视频模型对跌倒等短暂安全事件的时间稀疏性;以及系统性年龄歧视和价值偏差。结论认为当前架构不足以自主完成老年精神科评估,近期应限于低风险辅助功能并配合人工监督。
为什么推荐给您:系统梳理大语言模型在老年精神科评估中的失败模式,为安全应用提供新框架。
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
BACKGROUND: Large language models (LLMs) show strong performance on medical examinations, yet their reliability in real-world psychogeriatric assessment is uncertain, where delirium, dementia, and late-life depression often coexist amid frailty, polypharmacy, and long clinical histories. We aimed to map empirically evaluated LLM failure modes most relevant to geriatric psychiatric assessment and risk management.
METHODS: We conducted a systematic scoping review of empirical evaluations published between January 1, 2023, and December 31, 2025. We searched biomedical and technical sources and included studies that tested LLMs on clinically relevant tasks, including diagnostic reasoning, longitudinal history integration, safety and medication reasoning, and bias-related outcomes. Evidence was charted and synthesized qualitatively to derive a pragmatic taxonomy of recurrent failure patterns.
RESULTS: Forty-seven studies met inclusion criteria. Four convergent failure modes emerged: (1) diagnostic instability in longitudinal reasoning, including degraded retrieval within long contexts and loss of mid-history cues; (2) adversarial vulnerability, with elevated hallucination under misleading inputs and flawed rationales despite correct final answers; (3) multimodal temporal sparsity in video-based models that can omit brief safety-critical events such as falls; and (4) systemic ageism and value misalignment that can distort clinical narratives, risk estimation, and care recommendations.
CONCLUSIONS: Current LLM architectures remain insufficient for autonomous geriatric psychiatric assessment. Near-term use should be limited to low-risk support functions with human oversight, structured cognitive forcing safeguards, and evaluation protocols that stress-test longitudinal complexity, adversarial conditions, multimodal safety events, and bias.