Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require.
arXiv:2601. 22025v2 Announce Type: replace-cross Abstract: Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically variable, and sensitive to prompt and model changes.
By Daniel Commey
arXiv:2607. 24991v1 Announce Type: cross Abstract: Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs).
By Barbara Kitchenham, Sebasti\'an Pizard, Lech Madeyski, Ronnie de Souza Santos, Martin Shepperd, David Budgen
arXiv:2606. 28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates the complete systematic review and meta-analysis (SR/MA) workflow -- from literature search through statistical analysis, manuscript generation, and quality assurance -- with mandatory human oversight at critical decision points.
By Hsieh-Ting Lin, Jiunn-Tyng Yeh
SciLitBench is a multi-stage benchmark for evaluating large language models (LLMs) in systematic literature reviews, covering title and abstract screening, full-text screening, and schema-guided data extraction across 42,981 records and 888 included papers. The study shows that explicit inclusion/exclusion criteria boost title and abstract screening performance by 28.8% and researcher-authored rationales improve full-text screening by 15%. Data extraction performance varies widely, with high accuracy for publication year but low overlap for computational approaches, and even the best models recover only a fraction of annotated evidence and limitations.
By Miguel Zabaleta, Baihan Lin
arXiv:2606. 17588v1 Announce Type: cross Abstract: Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting mixed accuracy.
By Mika M\"antyl\"a, Patricia Matsubara, Katia Romero Felizardo, Miikka Kuutila, Marco Gerosa, Savio de Sousa Sampaio, Tayana Conte, Igor Steinmacher