arXiv AI

Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

The paper introduces the concept of an epistemic Sybil problem in multi‑agent AI systems, where multiple agents may produce seemingly independent reports that actually stem from the same underlying evidence. It formalizes this issue using information‑theoretic measures and demonstrates through large‑scale experiments that naive aggregation of replicated reports can severely degrade inference accuracy unless the system accounts for shared evidence ancestry and correlated extraction errors. The study shows that aggregators that track evidential dependence rather than merely report multiplicity or similarity achieve better calibration and inference performance.

arXiv AI
6d ago

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

The paper introduces the Correlated Promotion Benchmark (CPB) to evaluate how agents decide whether to admit claims into shared memory, addressing the risk of repeating false claims. CPB offers two modes: CPB-Static, a frozen test set with fixed gold actions, and CPB-Live, which runs multi‑agent teams and tracks source lineage. Experiments across eight admission policies and four agent families show that deduplication reduces false claims but also discards true ones, while gating on declared source type most effectively limits false adoption.

By Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv AI
Sep 3

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.

By Peiying Zhu, Sidi Chang
arXiv AI
Sep 10

Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification

The paper presents a framework for causal attribution in agentic AI systems, outlining estimators and conditions where they fail. It distinguishes between marginal total effects and common‑random‑number total effects, introduces a natural direct effect under pinned downstreams, and derives a coupling method to keep direct effects estimable. The authors also propose a traceability specification to meet upcoming regulatory requirements for high‑risk AI systems.

By Ajay Pravin Mahale (Hochschule Trier)
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin