arXiv:2608. 03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.
By Molood Arman
arXiv:2608. 03722v2 Announce Type: replace Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.
By Molood Arman
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion.
The paper introduces the Correlated Promotion Benchmark (CPB) to evaluate how agents decide whether to admit claims into shared memory, addressing the risk of repeating false claims. CPB offers two modes: CPB-Static, a frozen test set with fixed gold actions, and CPB-Live, which runs multi‑agent teams and tracks source lineage. Experiments across eight admission policies and four agent families show that deduplication reduces false claims but also discards true ones, while gating on declared source type most effectively limits false adoption.
By Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai
TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.
By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.
By Peiying Zhu, Sidi Chang
arXiv:2606. 18037v1 Announce Type: new Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools.
By Ander Alvarez, Santhiya Rajan, Samuel Mugel, Rom\'an Or\'us
The paper presents a framework for causal attribution in agentic AI systems, outlining estimators and conditions where they fail. It distinguishes between marginal total effects and common‑random‑number total effects, introduces a natural direct effect under pinned downstreams, and derives a coupling method to keep direct effects estimable. The authors also propose a traceability specification to meet upcoming regulatory requirements for high‑risk AI systems.
By Ajay Pravin Mahale (Hochschule Trier)
arXiv:2606. 10794v3 Announce Type: replace Abstract: Existing black-box LLM provenance methods achieve comparability by querying every candidate model with the same diagnostic prompts.
By Jiaxu Liu, Sunnan Mu, Dong Huang, Liuyin Wang, Jing Shao, Jie Zhang
arXiv:2606. 00005v1 Announce Type: new Abstract: We present the Consilium Protocol, a Byzantine Fault Tolerance-derived architecture for structured multi-model AI deliberation that treats inter-model disagreement as epistemic signal rather than error.
By VD Doske
The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.
By Qing Ye, Meng-Hsuan Lin
arXiv:2607. 09689v3 Announce Type: replace Abstract: Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged.
By Yossi Eliaz