arXiv AI By Barbara Kitchenham, Sebasti\'an Pizard, Lech Madeyski, Ronnie de Souza Santos, Martin Shepperd, David Budgen

Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

Read the original on arXiv AI →

arXiv:2607. 24991v1 Announce Type: cross Abstract: Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 14

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.

By Miguel Zabaleta, Baihan Lin
arXiv Computation and Language
Sep 22

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

The paper introduces CoSLR, a Human‑AI collaborative system for systematic literature reviews that incorporates mandatory human checkpoints within a three‑phase pipeline using large language models and Retrieval‑Augmented Generation. In a survey of 63 participants, 42.9 % rated the system’s usability highly, yet 34.9 % indicated they would trust AI‑generated summaries without further human verification after brief interaction. The study highlights that effective human oversight in AI‑assisted literature reviews depends on users’ willingness to engage with the checkpoints, underscoring a calibration issue that interface design must directly address.

By MD Aidul Islam, Malik Abdul Sami, Muhammad Waseem, Zeeshan Rasheed, Kai-kristian Kemell, Zheying Zhang, Pekka Abrahamsson
arXiv AI
Jun 30

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

arXiv:2606. 28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates the complete systematic review and meta-analysis (SR/MA) workflow -- from literature search through statistical analysis, manuscript generation, and quality assurance -- with mandatory human oversight at critical decision points.

By Hsieh-Ting Lin, Jiunn-Tyng Yeh
arXiv AI
Jul 10

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

arXiv:2607. 07980v1 Announce Type: cross Abstract: Coding agents now author entire pull requests, and practitioners sharply disagree about what this does to code review: whether it becomes the bottleneck, whether human review is still necessary, and whether it quietly erodes the understanding that it once built.

By Shyam Agarwal, Courtney Miller, Christian K\"astner, Bogdan Vasilescu