arXiv AI

Automated reproducibility assessments in the social and behavioral sciences using large language models

arXiv:2606. 13670v1 Announce Type: new Abstract: Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered.

arXiv AI
Sep 10

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

SciLitBench is a multi-stage benchmark for evaluating large language models (LLMs) in systematic literature reviews, covering title and abstract screening, full-text screening, and schema-guided data extraction across 42,981 records and 888 included papers. The study shows that explicit inclusion/exclusion criteria boost title and abstract screening performance by 28.8% and researcher-authored rationales improve full-text screening by 15%. Data extraction performance varies widely, with high accuracy for publication year but low overlap for computational approaches, and even the best models recover only a fraction of annotated evidence and limitations.

By Miguel Zabaleta, Baihan Lin
arXiv AI
Sep 16

Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits

The paper examines whether removing declared language fields from de‑identified résumés eliminates demographic leakage in large language models. By keeping language attributes identical and varying only unstructured prose across five ethnocultural groups and three cue‑salience levels, the authors find that non‑language text still allows target‑group recovery (average 0.757, reaching 1.000 under high salience). They also show that evaluation design—such as allowing or forbidding ties—dramatically affects LLM‑as‑a‑judge outcomes, underscoring the importance of evaluation protocol in bias audits.

By Qiangju Chen, Yang Xiao
arXiv Computation and Language
Sep 22

LLJ Cards: Best practices for the Use of LLMs as Judges

arXiv:2609.24516v1 Announce Type: new Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...

By Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung
arXiv AI
Sep 2

Medical Causal Hypothesis Verification with Large Language Models

The paper "Medical Causal Hypothesis Verification with Large Language Models" reports a small-scale study evaluating eight LLMs on 17 medical causal hypotheses. The authors introduce an evaluation framework and annotate 1,067 evidence points across six criteria, using nine metrics to assess performance. Results show that while LLMs have strong recall, they frequently fail to provide valid scientific articles, evidence, or reject unsupported hypotheses, revealing a critical limitation for their use in healthcare.

By Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva