UTFacultiesEEMCSDisciplines & departmentsSCSEducationAssignmentsFinished AssignmentsFinished Master AssignmentsBenchmarking AI-Based Tools Across PRISMA Phases: Toward Standardized Evaluation and AI-Assisted Workflows for Systematic Literature Reviews

Benchmarking AI-Based Tools Across PRISMA Phases: Toward Standardized Evaluation and AI-Assisted Workflows for Systematic Literature Reviews

MASTER Assignment

Benchmarking AI-Based Tools Across PRISMA Phases: Toward Standardized Evaluation and AI-Assisted Workflows for Systematic Literature Reviews

Type : Master M-BIT

Period: November 2025 - May, 2026

Student: Motika, H. (Haris, Student M-BIT)

Date Final project: May 8, 2026

Thesis

Supervisors:

Abstract:

Background: Systematic literature reviews (SLRs) are essential for synthesizing scientific evidence, but the rapidly increasing volume of academic publications has made them increasingly time-consuming. In particular, screening and data extraction phases often require the manual review of thousands of studies.

Objective: This study aims to evaluate the performance of AI-based tools, specifically large language models (LLMs), across multiple phases of the systematic literature review process and to see if these tools can be integrated into an AI-assisted automation pipeline.

Methods: A Python-based evaluation pipeline was developed to benchmark the GPT-5-Nano model across three reconstructed datasets who are defined as ”golden datasets”. The model was evaluated on Title & Abstract screening, Full-Text screening and Data Extraction using standardized metrics derived from the confusion matrices and existing literature.

Results: During Title and Abstract screening, recall ranged from 71.6% to 100.0%, while achieving workload reductions between 90.2% and 97.3%. Full-Text screening achieved recall values between 83.3% and 88.5%, though workload savings were lower. Data extraction performance varied substantially across datasets, with recall ranging from 45.5% to 94.4%, particularly affected by document structure and multi-value variables. Important to note that 100.0% recall was achieved on a dataset with relatively small sample size (n=509).

Conclusion: AI-based tools show strong potential for supporting systematic literature review workflows, particularly in early screening phases. However, variability in performance across tasks suggests that
fully automated pipelines are not yet sufficiently reliable. A human-in-the loop AI-assisted pipeline is proposed and should balance workload savings while maintaining methodological reliability.