arXiv AI

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

SciLitBench is a multi-stage benchmark for evaluating large language models (LLMs) in systematic literature reviews, covering title and abstract screening, full-text screening, and schema-guided data extraction across 42,981 records and 888 included papers. The study shows that explicit inclusion/exclusion criteria boost title and abstract screening performance by 28.8% and researcher-authored rationales improve full-text screening by 15%. Data extraction performance varies widely, with high accuracy for publication year but low overlap for computational approaches, and even the best models recover only a fraction of annotated evidence and limitations.

arXiv Computation and Language
6d ago

A Benchmark Framework for Screening Automation in Systematic Reviews

The paper introduces SRBench, a benchmark dataset comprising 45,064 labeled entries from 32 curated secondary studies, designed to evaluate large language model performance in systematic review screening while addressing class imbalance. It also presents PromptSR, a tool that facilitates prompt experimentation, experiment management, and result analysis for LLM-based screening. A use case demonstrates the practical application of both SRBench and PromptSR.

By Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani
arXiv AI
Jun 19

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

arXiv:2606. 19345v1 Announce Type: cross Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent.

By Zhyar Rzgar K. Rostam, M\'arta P\'entek, J\'anos Tibor Czere, Zsombor Zrubka, L\'aszl\'o Gul\'acsi, G\'abor Kert\'esz
arXiv Computation and Language
Sep 14

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.

By Miguel Zabaleta, Baihan Lin
arXiv AI
Aug 20

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.

By Valentin Romanov, Monique Bax, Steven Niederer
arXiv AI
Jun 30

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

arXiv:2606. 28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates the complete systematic review and meta-analysis (SR/MA) workflow -- from literature search through statistical analysis, manuscript generation, and quality assurance -- with mandatory human oversight at critical decision points.

By Hsieh-Ting Lin, Jiunn-Tyng Yeh
arXiv AI
Aug 28

Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

The paper presents a pipeline that uses Large Language Models (LLMs) to extract information from 536 peer‑reviewed agent‑based modeling papers for systematic literature reviews (SLRs). GPT‑4.1 achieves about 77.95% paper‑level accuracy, while GPT‑5.0 reaches 81.67%. Field‑level accuracy varies widely, and the study notes that agreement between LLMs can signal output quality, with low agreement indicating hallucinations and high agreement with low accuracy suggesting noise in the human dataset.

By Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak
arXiv AI
Sep 24

TiAb Review Plugin: A Browser-Based Tool for AI-Assisted Study Selection in Systematic Reviews

arXiv:2604.08602v2 Announce Type: replace-cross Abstract: Server-based screening tools impose subscription costs, while open-source alternatives require coding skills, and full-text screening has rem...

By Yuki Kataoka, Masahiro Banno, Michihito Kyo, Shuri Nakao, Tomoo Sato, Shunsuke Taito, Tomohiro Takayama, Takahiro Tsuge, Yasushi Tsujimoto, Ryuhei So, Toshi A. Furukawa
Hugging Face Trending Papers
Aug 19

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.

arXiv AI
Sep 2

Medical Causal Hypothesis Verification with Large Language Models

The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.

By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva
arXiv AI
Aug 28

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

The study evaluates literature reviews produced by large language models (LLMs) using short and long context windows, assessing their quality across 15 dimensions. Results show that while larger context windows allow LLMs to incorporate more information and maintain coherence, they also increase repetition, omission of key works, and a tendency toward descriptive rather than synthetic content. Human oversight remains essential for meeting academic publishing standards, and the authors suggest future work should blend human expertise with AI to mitigate these limitations.

By Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby