arXiv AI By Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Clayton Suguio Hida

Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews

Read the original on arXiv AI →

arXiv:2608. 14737v1 Announce Type: cross Abstract: This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
6d ago

A Benchmark Framework for Screening Automation in Systematic Reviews

The paper introduces SRBench, a benchmark dataset comprising 45,064 labeled entries from 32 curated secondary studies, designed to evaluate large language model performance in systematic review screening while addressing class imbalance. It also presents PromptSR, a tool that facilitates prompt experimentation, experiment management, and result analysis for LLM-based screening. A use case demonstrates the practical application of both SRBench and PromptSR.

By Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani
arXiv AI
Aug 28

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

The study compared human and large language model (LLM) workflows for title‑and‑abstract screening in a complex scoping review. Human reviewers and two GPT‑5.4 file‑batch runs retained 42.2‑45.0% of records with 82.3‑82.9% recall, while Gemini 3.1 achieved the highest recall (83.9%) but retained 56.7% of records. Identical GPT‑5.4 runs showed 91.7% agreement yet differed on 94 records, including 29 verified eligible ones.

By Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig
arXiv Computation and Language
Sep 14

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.

By Miguel Zabaleta, Baihan Lin
arXiv AI
Jun 19

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

arXiv:2606. 19345v1 Announce Type: cross Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent.

By Zhyar Rzgar K. Rostam, M\'arta P\'entek, J\'anos Tibor Czere, Zsombor Zrubka, L\'aszl\'o Gul\'acsi, G\'abor Kert\'esz
arXiv AI
Aug 20

When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators

The paper evaluates large language models (LLMs) as data quality annotators on two e-commerce tasks: entity matching and brand mislabeling. In entity matching, a simple rule-based baseline matched the LLM’s zero-shot performance (F1≈0.95), and a few-shot prompt actually lowered performance, highlighting the risk of small-sample prompt tuning. For brand mislabeling, the LLM outperformed a naive rule baseline (F1 0.833 vs 0.721) by leveraging background knowledge, and demonstrated high consistency across repeated runs (99.7% agreement).

By Praphulla Lal Shrestha