医学伦理研究助手

利用大语言模型构建高内容效度的可解释文本模型词典:自杀风险词典

Using large language models to create lexicons for interpretable text models with high content validity: The suicide risk lexicon

J Psychopathol Clin Sci · 2026 年 9 月 24 日 · Daniel M Low, Osiris Rankin, Daniel D L Coppersmith 等 6 人

建模 / 计算研究
在聊天里讨论
一分钟了解
用GPT-4 Turbo自动生成自杀风险词典,在危机咨询对话中预测风险并媲美部分黑箱模型。

研究者常需从问卷、访谈、社交媒体和电子病历文本中测量焦虑、孤独等概念,大语言模型虽分类效果最好,但受成本、隐私和算力限制并非人人可用,轻量可解释的词典方法因此仍有价值,但人工构建词典费时费力。本研究用GPT-4 Turbo自动为49个已知自杀意念与行为风险因素生成词典,即“自杀风险词典”,内容效度较高,在危机咨询对话中能准确预测风险。经临床医生评分验证后,该词典略优于语言探究与词数计数词典,与部分黑箱深度学习模型表现相当。研究还发现,在该真实咨询场景中,主动自杀意念和直接自伤比被动自杀意念和抑郁情绪更能提示迫近风险。作者发布了Python工具包construct-tracker以便其他领域构建词典。

为什么推荐给您:用大语言模型自动构建可解释词典并验证其危机预测价值,属方法学新进展,需注意隐私。

讲解深度:

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

摘要Abstract

摘要第 1 段问这一段

Researchers often want to measure a variety of constructs such as anxiety, discrimination, or loneliness in text data from surveys, interviews, social media, and electronic health records. Using large language models (LLMs)-although optimal for text classification-remain infeasible for some researchers due to concerns around computational expertise, cost, privacy, and compute requirements. Therefore, some researchers prefer to use lightweight models for large data sets or interpretable models to avoid mistakes in high-stakes scenarios. Lexicons offer simple baselines to LLMs by searching for relevant phrases and can be used together with LLMs to guarantee capturing specific keywords. However, building new lexicons is resource intensive. In this study, we found that GPT-4 Turbo was able to automatically create a lexicon for 49 known risk factors for suicidal thoughts and behaviors, which we release as the Suicide Risk Lexicon. Generating a lexicon with LLMs quickly measures most constructs relevant for suicide risk detection, resulting in high content validity. This lexicon was able to accurately predict risk in crisis counseling conversations. After validating the lexicon with clinician ratings, the lexicon modestly outperformed the linguistic inquiry and word count lexicon, which has low content validity for mental illness, and performed similarly to some black-box deep learning models. We discovered that active suicidal ideation and direct self-injury were stronger indicators of imminent risk than passive suicidal ideation and depressed mood in this ecological setting of crisis counseling conversations. To simplify creating new lexicons for other research domains, we introduce a Python package, construct-tracker. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

从这篇论文记下的摘录
在“讲解”“原文”里选中文字,会出现“记到笔记”按钮(电脑上在文字旁边,手机上在屏幕最下面);记下的内容会按笔记本整理,也会列在这里。
讲解或动画有问题?告诉我: