arXiv AI By Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma, Frauke Kreuter, Markus Weinmann, Stefan Feuerriegel

Automated reproducibility assessments in the social and behavioral sciences using large language models

Read the original on arXiv AI →

arXiv:2606. 13670v1 Announce Type: new Abstract: Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

SciLitBench is a multi-stage benchmark for evaluating large language models (LLMs) in systematic literature reviews, covering title and abstract screening, full-text screening, and schema-guided data extraction across 42,981 records and 888 included papers. The study shows that explicit inclusion/exclusion criteria boost title and abstract screening performance by 28.8% and researcher-authored rationales improve full-text screening by 15%. Data extraction performance varies widely, with high accuracy for publication year but low overlap for computational approaches, and even the best models recover only a fraction of annotated evidence and limitations.

By Miguel Zabaleta, Baihan Lin