arrow
Return

Automating the data extraction process for systematic reviews using GPT-4o and o3

delete2025-09-01
delete1
delete
OA
AI
Y
Yuki Kataoka *
T
T. Takayama
Y
Yoshimura, Keisuke
R
Ryuhei So
Y
Yasushi Tsujimoto
Y
Yosuke Yamagishi
S
Shiro Takagi
Y
Yuki Furukawa
M
Masatsugu Sakata
Đ
Đorđe Bašić
A
Andrea Cipriani
P
Pim Cuijpers
Z
Zhu, Tengfei
M
Mathias Harrer
S
Stefan Leucht
A
Ava Homiar
E
Edoardo G. Ostinelli
C
Clara Miguel
A
Alessandro Rodolico
T
Toshi A. Furukawa
DOI:10.1017/rsm.2025.10030delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Large language models have shown promise for automating data extraction (DE) in systematic reviews (SRs), but most existing approaches require manual interaction. We developed an open-source system using GPT-4o to automatically extract data with no human intervention during the extraction process. We developed the system on a dataset of 290 randomized controlled trials (RCTs) from a published SR about cognitive behavioral therapy for insomnia. We evaluated the system on two other datasets: 5 RCTs from an updated search for the same review and 10 RCTs used in a separate published study that had also evaluated automated DE. We developed the best approach across all variables in the development dataset using GPT-4o. The performance in the updated-search dataset using o3 was 74.9% sensitivity, 76.7% specificity, 75.7 precision, 93.5% variable detection comprehensiveness, and 75.3% accuracy. In both datasets, accuracy was higher for string variables (e.g., country, study design, drug names, and outcome definitions) compared with numeric variables. In the third external validation dataset, GPT-4o showed a lower performance with a mean accuracy of 84.4% compared with the previous study. However, by adjusting our DE method, while maintaining the same prompting technique, we achieved a mean accuracy of 96.3%, which was comparable to the previous manual extraction study. Our system shows potential for assisting the DE of string variables alongside a human reviewer. However, it cannot yet replace humans for numeric DE. Further evaluation across diverse review contexts is needed to establish broader applicability.
Keywords:
data extraction automation
GPT-4o
large language models
o3
systematic reviews
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Research Synthesis Methods cover
Research Synthesis Methods
IF:
6.1
Papers:
826
Citations:
8.7K

Organization

T
tohoku university
Scholars:
4.3W
Papers: 3.6W
Citations: 31
V
Vrije Universiteit Amsterdam
Scholars:
4.2W
Papers: 3.7W
Citations: 3.7W
O
Oxford Health NHS Foundation Trust
Scholars:
155
Papers: 95
Citations: 1.1K
U
university of tokyo
Scholars:
6.3K
Papers: 2.5K
Citations: 1
U
university of munich
Scholars:
1.4K
Papers: 500
Citations: 0
K
kyoto university
Scholars:
7.6K
Papers: 3.0K
Citations: 0
researcher View more organizations