大语言模型用于肝细胞癌术前微血管侵犯预测:与放射科医师的多中心比较及治疗结局
Large Language Models for Preoperative Microvascular Invasion Prediction in Hepatocellular Carcinoma: A Multicenter Comparison with Radiologists and Treatment Outcomes
肝细胞癌手术前能否判断肿瘤已侵犯微小血管(微血管侵犯),关系到复发风险和治疗方案选择,但靠影像判断一直很难。研究回顾了602例术前做过磁共振、随后手术切除的肝癌患者,让GPT-4o和DeepSeek-R1根据影像报告文字做预测,并与6位不同年资的放射科医师比较,以病理结果为金标准。结果DeepSeek-R1准确率最高(81.2%),高于所有参与比较的医师;GPT-4o与资深医师相当,且在不同肿瘤大小下表现较稳定。模型判断为微血管侵犯阳性的患者复发更快,且解剖性切除比非解剖性切除效果更好,提示这一方法可能帮助分层治疗。属于回顾性研究,尚需前瞻性验证。
为什么推荐给您:大语言模型用于术前影像风险分层并超越多数医师,属方法学重要进展,但为回顾性研究。
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…
已等待 0 秒
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。
摘要Abstract
BACKGROUND: Preoperative prediction of microvascular invasion (MVI) in hepatocellular carcinoma (HCC) is critical for prognosis but challenging. This study evaluated the performance of large language models (LLMs) for MVI prediction compared with radiologists with varying levels of experience and explored the association of model-predicted MVI with recurrence and treatment outcomes.
MATERIALS AND METHODS: In this retrospective multicenter study, 602 patients with pathologically confirmed HCC who underwent preoperative MRI and surgical resection were included. GPT-4o and DeepSeek-R1, prompted in English and Chinese using radiology report narratives and predefined MRI features, were compared with six radiologists using histopathology as the reference standard. An independent cohort of 135 patients who underwent radiofrequency ablation (RFA) was included for treatment-outcome analyses. Prediction performance was assessed using diagnostic accuracy and the area under the curve (AUC).
RESULTS: GPT-4o (English input) achieved 75.6% accuracy, outperforming residents (59.3%, P < 0.001) and attendings (66.7%, P < 0.001), and was comparable to senior radiologists (74.2%, P = 1.000). DeepSeek-R1 (Chinese input) achieved the highest accuracy (81.2%), outperforming all radiologists (P < 0.05). Residents and attendings showed high sensitivity but relatively low specificity, whereas senior radiologists demonstrated a more balanced performance; while DeepSeek-R1 showed better specificity (up to 94.2%, P < 0.001). LLMs performance were relatively stable across tumor sizes, whereas accuracy among less-experienced radiologists improved with tumor size. AUCs for LLMs (GPT-4o: 75.4%; DeepSeek-R1 Chinese: 80.1%) were superior or comparable to those of radiologists (61.8%-74.5%). Stacking ensemble-predicted MVI-positive status was associated with shorter time-to-recurrence (TTR) in both small (≤3 cm, P = 0.039) and large (>3 cm, P = 0.001) tumors, and with better outcomes following anatomical rather than non-anatomical resection (≤3 cm, P = 0.021; >3 cm, P < 0.001). After propensity score matching, surgical resection was associated with longer TTR than RFA in the model-predicted MVI-positive subgroup, but not in the MVI-negative subgroup.
CONCLUSION: LLMs, particularly DeepSeek-R1 and GPT-4o, demonstrated diagnostic performance for preoperative MVI prediction comparable to or better than that of senior radiologists, with relatively stable performance across tumor sizes and potential value for recurrence-risk stratification and treatment planning.