医学伦理研究助手

基于集体智能与模仿学习的人工智能诊断决策支持系统用于老年人基层医疗

Artificial intelligence-based diagnostic decision support for primary care of older adults using collective intelligence and imitation learning

J Am Med Inform Assoc · 2026 年 9 月 24 日 · Christopher D Streiffer, Matthew J Press, Nicholas S Bishop 等 9 人

建模 / 计算研究
在聊天里讨论
一分钟了解
用模仿学习训练AI,为老年基层患者生成诊断和检查建议。

老年人基层医疗中诊断错误风险较高,但现有诊断决策支持系统依赖难以获得的金标准标签、诊断范围窄、也无法反映真实世界的不确定性。研究者利用模仿学习和临床医生集体智能,开发并验证了一个深度学习诊断决策支持系统:用70多万次65岁以上老年人基层就诊的结构化和非结构化电子病历数据,训练多标签神经网络模仿临床医生的决策,生成诊断和医嘱建议,覆盖669种诊断和1000种医嘱。在时间上独立保留的测试集上,诊断建议表现优异(微观c统计量0.995,宏观c统计量0.896),但医嘱预测明显较弱。随后用随机、盲法的等效性研究比较模型建议与真实临床决策,结果显示诊断不一致率在模型与真实实践之间统计学等效,而医嘱不一致率不等效。该研究提出了一条无需专家标注即可生成贴近真实诊疗建议的新路径,但临床验证仍是初步的。

为什么推荐给您:用模仿学习和集体医生智能构建诊断支持系统,方法新颖且样本庞大,但临床验证仍属初步。

讲解深度:

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

摘要Abstract

摘要第 1 段问这一段

OBJECTIVE: Older adults face disproportionate risk of diagnostic errors in primary care, yet effective diagnostic decision support systems (DDSSs) remain lacking. Existing artificial intelligence (AI)-based DDSSs rely on gold-standard labels infeasible to obtain, have narrow diagnostic scope, and fail to reflect real-world uncertainty. We developed and validated a deep learning DDSS using imitation learning and collective clinician intelligence to generate diagnostic and order recommendations.

摘要第 2 段问这一段

MATERIALS AND METHODS: We trained a multi-label neural network on structured and unstructured electronic health record data from primary care encounters of adults ≥65 years old to imitate clinician decisions. Models were evaluated using discrimination, calibration, threshold-based, and composite metrics on a temporally held-out test set. Clinical validity was assessed using a randomized, blinded equivalence study comparing model-generated recommendations with observed clinician decisions.

摘要第 3 段问这一段

RESULTS: The study included 707 598 primary care encounters, generating recommendations across 669 diagnoses and 1000 orders. The general model demonstrated excellent performance (micro c-statistic 0.995 [95% CI 0.995-0.996], macro c-statistic 0.896 [0.888-0.897], micro F1-score 0.904 [0.903-0.905], macro F1-score 0.384 [0.381-0.385], integrated calibration index (ICI) 0.0025 [0.0024-0.0026]). Order prediction was lower (micro c-statistic 0.812 [0.811-0.814], macro c-statistic 0.786 [0.777-0.786], micro F1-score 0.305 [0.303-0.307], macro F1-score 0.0286 [0.0277-0.0294], ICI 0.0145 [0.0141-0.0150]). In clinician-validation, diagnostic disagreements were statistically equivalent between model recommendations and observed practice while order disagreements were not.

摘要第 4 段问这一段

CONCLUSION: Imitation learning and collective clinician intelligence provide a feasible framework for generating diagnostic and ordering recommendations that reflect real-world practice without expert-adjudicated training labels. Randomized, blinded clinician evaluation pragmatically establishes preliminary clinical validity beyond traditional retrospective metrics.

从这篇论文记下的摘录
在“讲解”“原文”里选中文字,会出现“记到笔记”按钮(电脑上在文字旁边,手机上在屏幕最下面);记下的内容会按笔记本整理,也会列在这里。
讲解或动画有问题?告诉我: