The paper introduces SRBench, a benchmark dataset comprising 45,064 labeled entries from 32 curated secondary studies, designed to evaluate large language model performance in systematic review screening while addressing class imbalance. It also presents PromptSR, a tool that facilitates prompt experimentation, experiment management, and result analysis for LLM-based screening. A use case demonstrates the practical application of both SRBench and PromptSR.
By Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani
arXiv:2606. 19345v1 Announce Type: cross Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent.
By Zhyar Rzgar K. Rostam, M\'arta P\'entek, J\'anos Tibor Czere, Zsombor Zrubka, L\'aszl\'o Gul\'acsi, G\'abor Kert\'esz
The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.
By Miguel Zabaleta, Baihan Lin
The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
By Valentin Romanov, Monique Bax, Steven Niederer
arXiv:2606. 28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates the complete systematic review and meta-analysis (SR/MA) workflow -- from literature search through statistical analysis, manuscript generation, and quality assurance -- with mandatory human oversight at critical decision points.
By Hsieh-Ting Lin, Jiunn-Tyng Yeh
The paper presents a pipeline that uses Large Language Models (LLMs) to extract information from 536 peer‑reviewed agent‑based modeling papers for systematic literature reviews (SLRs). GPT‑4.1 achieves about 77.95% paper‑level accuracy, while GPT‑5.0 reaches 81.67%. Field‑level accuracy varies widely, and the study notes that agreement between LLMs can signal output quality, with low agreement indicating hallucinations and high agreement with low accuracy suggesting noise in the human dataset.
By Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak