arXiv AI

Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility

The paper argues that empirical results in Machine Learning are often difficult to reproduce due to limited availability of code and supporting materials, which hampers research progress. It analyzes and quantifies these challenges and proposes concrete measures to enhance the verifiability of results, even if full reproducibility cannot be guaranteed. The authors provide their code and supporting resources on GitHub for reference.

arXiv Machine Learning
Sep 16

OPEN-1B: A Fully Auditable Training Run

The paper introduces Open-1B, a language model trained under a new fully auditable regime that ensures every training operation is reproducible on heterogeneous commodity hardware with bitwise certainty. By enforcing a fixed order on sources of nondeterminism—GPU reductions, data batch ordering, and inter/intra-node communication—the authors enable auditors to replay and verify individual training steps on a single machine. The release includes the full pretraining dataset, all intermediate checkpoints, the training codebase, and an audit harness for step-by-step verification.

By John Donaghy, Brian Wilcox, O\u{g}uzhan Ersoy, Shikhar Rastogi, Adam St Arnaud, Alexey Titov, Jordan Greenberg, Ben Fielding, Harry Grieve
arXiv AI
Aug 28

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

The paper "Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research" argues that large language model agents must faithfully implement reference methods, design experiments that truly test claims, and provide supporting evidence. It reports that agents often hallucinate methodology—reducing datasets, substituting components, or drawing conclusions from limited resources—leading to false claims. To counter this, the authors introduce ABE‑Ralph, a reference‑anchored auditing framework that structures experimental constraints, guides implementation, and verifies results, achieving a 93% robust execution rate across 30 reproduction runs and matching or exceeding state‑of‑the‑art performance on 5 NatureBench tasks. "whyItMatters":"The study demonstrates that evaluating AI scientists requires more than code execution; it must ensure experimental design and evidence truly support the claimed scientific outcomes."

By Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan
arXiv AI
Aug 24

Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

The paper argues that prediction‑based certifications—such as accuracy, calibration, and conformal coverage—are insufficient to guarantee trustworthy AI. It proves a separation theorem showing that a model can appear reliable under all prediction‑side certificates yet differ arbitrarily in explanation fidelity and deployment behaviour. The authors propose a competence envelope framework that combines both prediction and explanation certification to detect such hidden failures.

By Nataliya Shakhovska, Ivan Izonin, Stergios-Aristoteles Mitoulis
arXiv AI
Jun 16

Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs

arXiv:2606. 15258v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet we still lack a scalable and reproducible way to measure step-level reasoning in long proofs across diverse sources.

By Jierui Zhang, Siyuan Tan, Xinhang Li, Longzhuangzhi Lin, Dailin Li, Chengfeng Gu, Xinping Li, Yaxian Hao, Shengjia Liang, Yuxiang Ren, Wenhao Liu