论文 · 建模 / 计算研究
ProMem-agent:程序性记忆增强的大语言模型智能体用于临床病程推理
ProMem-agent: Procedural memory-augmented large language model agents for clinical trajectory reasoning
作者:Qianying He, Xuan Liu, Jingquan Liu, Bao Liu, Wenjian Liu
Int J Med Inform · 2026年9月17日 · He 等 5 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
BACKGROUND AND OBJECTIVE: Clinical large language model (LLM) agents can interpret current clinical context but have limited mechanisms for converting longitudinal experience into compact, reusable units. We developed ProMem-Agent, a procedural-memory framework that represents recurrent early intensive-care trajectories as provenance-linked observational patterns for retrospective mortality-risk estimation.
METHODS: Adult ICU stays with at least 24 hours of observable data were represented as six consecutive four-hour state-action-response intervals. Memory candidates were extracted exclusively from the MIMIC-IV ICU training cohort, linked to source events, consolidated by semantic clustering, and organized in a similarity graph. For each new patient, hybrid semantic and graph retrieval selected three memory cards. Comparators included conventional and longitudinal EHR models, direct and Chain-of-Thought LLM prompting, patient-level Case-RAG, token-matched Case-RAG, semantic-only Procedure-RAG, and first-24-hour SOFA as a clinically established severity reference. Evaluation included patient-level bootstrap testing, calibration analysis, external validation, retrieval and clustering sensitivity analyses, perturbation experiments, and blinded expert review.
RESULTS: The internal test cohort contained 6368 ICU stays with 11.9% mortality. ProMem-Agent achieved an F1-score of 0.608, AUC of 0.836, AUPRC of 0.481, and Brier score of 0.086. Relative to Procedure-RAG, the incremental differences were modest (AUC +0.013; AUPRC +0.029). Without memory reconstruction or external recalibration, AUC/AUPRC values were 0.823/0.402 on eICU, 0.831/0.429 on MIMIC-III, and 0.803/0.361 on HiRID. External calibration slopes were 0.88, 0.91, and 0.85, respectively, compared with 0.97 internally.
CONCLUSIONS: Procedural abstraction accounted for a larger share of the observed gain than graph propagation, while external miscalibration and dataset shift limited transportability of absolute risk. ProMem-Agent should therefore be interpreted as a retrospective research framework for studying reusable clinical-trajectory representations, not as a clinically deployable decision-support system. Prospective, site-specific, clinician-in-the-loop evaluation remains necessary.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。