arXiv Computation and Language

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.

arXiv AI
Jun 30

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

arXiv:2606. 28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates the complete systematic review and meta-analysis (SR/MA) workflow -- from literature search through statistical analysis, manuscript generation, and quality assurance -- with mandatory human oversight at critical decision points.

By Hsieh-Ting Lin, Jiunn-Tyng Yeh
arXiv AI
Sep 10

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

SciLitBench is a multi-stage benchmark for evaluating large language models (LLMs) in systematic literature reviews, covering title and abstract screening, full-text screening, and schema-guided data extraction across 42,981 records and 888 included papers. The study shows that explicit inclusion/exclusion criteria boost title and abstract screening performance by 28.8% and researcher-authored rationales improve full-text screening by 15%. Data extraction performance varies widely, with high accuracy for publication year but low overlap for computational approaches, and even the best models recover only a fraction of annotated evidence and limitations.

By Miguel Zabaleta, Baihan Lin
arXiv Computation and Language
Sep 24

EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine

EviStreams is a live, open‑source, no‑code web platform that enables systematic review teams to control AI‑assisted data extraction at three stages: program design, field specification, and extracted predictions. Reviewers use a form builder to define typed fields, run extraction on PDFs, inspect AI‑generated values with supporting passages, and perform blinded dual review with adjudication to produce an auditable consensus export. An evaluation across four clinical corpora and three model families shows that extraction quality depends more on field specification than on the model choice.

By Sai Karthik Kosuri, Ankita Shashikant Bhosale, Michael Glick, Alonso Carrasco-Labra, Chris Callison-Burch
arXiv AI
Jun 9

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.

By Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Max Lamparth, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
arXiv Computation and Language
6d ago

A Benchmark Framework for Screening Automation in Systematic Reviews

The paper introduces SRBench, a benchmark dataset comprising 45,064 labeled entries from 32 curated secondary studies, designed to evaluate large language model performance in systematic review screening while addressing class imbalance. It also presents PromptSR, a tool that facilitates prompt experimentation, experiment management, and result analysis for LLM-based screening. A use case demonstrates the practical application of both SRBench and PromptSR.

By Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani
arXiv AI
Jun 12

HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation

arXiv:2601. 19072v3 Announce Type: replace-cross Abstract: Large Language models (LLMs) have shown strong capabilities in code review automation, such as review comment generation, yet they suffer from hallucinations -- where the generated review comments are ungrounded in the actual code -- poses a significant challenge to the adoption of LLMs in code review workflows.

By Kla Tantithamthavorn, Hong Yi Lin, Patanamon Thongtanunam, Wachiraphan Charoenwet, Minwoo Jeong, Ming Wu