arXiv AI

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

arXiv:2607. 05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect.

arXiv AI
Aug 28

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

The paper "Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research" argues that large language model agents must faithfully implement reference methods, design experiments that truly test claims, and provide supporting evidence. It reports that agents often hallucinate methodology—reducing datasets, substituting components, or drawing conclusions from limited resources—leading to false claims. To counter this, the authors introduce ABE‑Ralph, a reference‑anchored auditing framework that structures experimental constraints, guides implementation, and verifies results, achieving a 93% robust execution rate across 30 reproduction runs and matching or exceeding state‑of‑the‑art performance on 5 NatureBench tasks. "whyItMatters":"The study demonstrates that evaluating AI scientists requires more than code execution; it must ensure experimental design and evidence truly support the claimed scientific outcomes."

By Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv AI
4d ago

Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy

The paper introduces a framework for evaluating large language model agents by attempting to end‑to‑end reproduce published astronomy studies, separating execution from verification and distinguishing computational failures from methodological ambiguities. Applying this to fourteen papers—one from The Astrophysical Journal and thirteen from Nature—revealed that eleven contained ambiguities that prevented a uniquely specified reproduction path. In a controlled case study, twelve different analysis paths produced distance estimates ranging from 2.16 to 3.53 kpc, with only one matching the published value of ~2.70 kpc, demonstrating that matching outcomes does not guarantee that the agent has reconstructed the underlying reasoning. whyItMatters":"The study shows that end‑to‑end reproduction can expose gaps in implicit scientific knowledge within AI systems, highlighting the need for better integration of causal relevance in LLM agents."

By Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang, Yangbin Xie, Xiaoyu Tang, Cong Sun