论文 · 综述
生成式人工智能的心理测量学应用:生命周期、风险与研究议程
Psychometric applications of generative artificial intelligence: Lifecycle, risks, and research agenda
作者:David Villarreal-Zegarra, Yscenia Paredes-Gonzales, Jackeline García-Serna
PLOS Ment Health · 2026年9月23日 · Villarreal-Zegarra 等 3 位作者
不需要生物学背景,多打比方
正在获取全文并生成讲解(拿不到全文就依据摘要)…
已等待 0 秒大约需要 10–20 秒
可以先看别的,做好了会自动出现在这里。
这篇还没有动画
动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要一两分钟。
摘要Abstract
Generative artificial intelligence (GenAI), including large language models and large multimodal models, is increasingly used to support psychological assessment and digital mental health measurement. This review proposes a lifecycle framework for evaluating these uses without weakening psychometric standards. The framework covers construct definition; item, prompt, indicator, or signal generation; evaluation of content, response processes, usability, and data quality; piloting and calibration; validation; fairness and measurement invariance; scoring and interpretation; documentation; and post-deployment monitoring. We argue that digital traces from smartphones, chatbots, ecological momentary assessment, wearables, social media, clinical notes, and multimodal systems should not be treated as psychometric measures by default. They should remain raw data, candidate indicators, or algorithmic scores until their construct interpretation and intended use are theoretically specified, technically verified, empirically calibrated, and validated with human data. GenAI may help generate candidate items, refine wording, classify open-text responses, extract structured information, support multimodal integration, and assist scoring under explicit rules. These uses may improve efficiency and scale, but they do not establish validity, objectivity, fairness, or clinical meaning. We distinguish two complementary facets: GenAI for psychometrics, in which models support measurement development and scoring, and psychometrics for GenAI, in which psychometric methods evaluate model behavior when LLMs are part of measurement workflows. Key risks include construct drift, face validity without structural validity, compressed variability in synthetic respondents, algorithmic bias, lack of invariance, prompt sensitivity, model drift, automation bias, data-security and privacy failures, and loss of subjectivity and disagreement. Responsible use requires human oversight, transparent reporting, validation with human data, fairness evaluation, secure data governance, and continuous monitoring. GenAI should augment, not replace, psychometric science.
还没有查过关联研究
我会去找这篇研究之前的基础工作、做类似事情的研究,以及之后引用它的研究,并说明每篇为什么相关。