arXiv:2608. 03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.
By Molood Arman
arXiv:2608. 03722v2 Announce Type: replace Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.
By Molood Arman
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion.
The paper introduces the Correlated Promotion Benchmark (CPB) to evaluate how agents decide whether to admit claims into shared memory, addressing the risk of repeating false claims. CPB offers two modes: CPB-Static, a frozen test set with fixed gold actions, and CPB-Live, which runs multi‑agent teams and tracks source lineage. Experiments across eight admission policies and four agent families show that deduplication reduces false claims but also discards true ones, while gating on declared source type most effectively limits false adoption.
By Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai
TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.
By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.
By Peiying Zhu, Sidi Chang