医学伦理研究助手
前沿
论文精读

论文

评估大语言模型使用Cochrane偏倚风险工具第2版进行偏倚风险评估的准确性:探索性可行性研究

J Med Internet Res · 2026年9月25日 · Lai 等 3 位作者

问这篇
一分钟了解要点用ChatGPT对29项随机试验做偏倚风险评估,准确率中等但一致性较高。结果ChatGPT各领域总体准确率约73%至76%,两次评估一致性较高,但在需要解读隐含叙述或行为细节的复杂情境中可靠性下降。

不需要生物学背景,多打比方

正在获取全文并生成讲解(拿不到全文就依据摘要),大约需要 30–60 秒…

已等待 0 秒

这篇还没有动画

动画会把研究的流程、作用机制和关键结果一步一步演示出来,每一步都标明出自原文哪里。制作大约需要 30–60 秒。

目前只拿到了摘要全文暂时拿不到(可能不是免费全文)。下面是论文摘要。

摘要Abstract

摘要第 1 段问这一段

BACKGROUND: Large language models (LLMs) have the potential to improve the efficiency of evidence synthesis, but their reliability in performing complex tasks such as risk-of-bias (ROB) assessment in randomized controlled trials (RCTs) remains unclear.

摘要第 2 段问这一段

OBJECTIVE: This study aimed to evaluate whether LLMs can reliably assess ROB in RCTs using version 2 of the Cochrane ROB tool for randomized trials (ROB 2).

摘要第 3 段问这一段

METHODS: This study was conducted between December 28, 2024, and February 28, 2025, in adherence to American Association for Public Opinion Research reporting guidelines. Twenty-nine RCTs were selected from published Cochrane systematic reviews across diverse medical fields. We developed a structured prompt engineering framework that transformed ROB 2 decision trees into logical rules for the LLM. Each RCT was independently evaluated twice by ChatGPT, with Cochrane review authors' assessments serving as the reference standard for comparison. The main outcomes were the accuracy and consistency of ROB 2 assessments at both the domain and trial levels, evaluated using accuracy, sensitivity, specificity, and F1-score. Consistency between the repeated assessments was quantified using the Cohen κ and prevalence-adjusted, bias-adjusted κ.

摘要第 4 段问这一段

RESULTS: The LLM demonstrated a moderate aggregate domain accuracy of 73.1% (95% CI 64.7%-81.5%) in the first assessment and 75.9% (95% CI 66.3%-85.4%) in the second assessment. Domain-averaged sensitivity decreased from 61.4% (95% CI 48.1%-74.7%) to 53.4% (95% CI 41.7%-65.0%), whereas domain-averaged specificity increased from 75.8% (95% CI 65.3%-86.3%) to 81.1% (95%CI 67.1%-95%), indicating a conservative tendency in identifying a high ROB. Domain-level accuracy ranged from 62.1% to 87.9%, with the lowest accuracy observed in domain 1 and the lowest F1-scores observed in domain 2. Consistency between repeated assessments was high, with a mean agreement of 89.0% (SD 7.5%), and Cohen κ values were 0.86, 0.39, 0.56, 0.84, and 0.85 in domains 1 to 5, respectively.

摘要第 5 段问这一段

CONCLUSIONS: In this exploratory study, ChatGPT demonstrated moderate accuracy and high consistency in assessing ROB in RCTs using the ROB 2. However, its reliability diminished in complex scenarios requiring interpretation of implicit narratives or behavioral nuance. These findings suggest that LLMs may support methodological evaluations in systematic reviews by acting as automated screeners to reduce reviewer burden, but current implementation still requires expert oversight, particularly for trials involving subjective outcomes or nonstandard reporting.

这篇对您:
讲解或动画有问题: