arXiv Computation and Language
Aug 27

Provenance Before Prose: Claim-Locked Reporting

The paper introduces claim‑locked reporting, a protocol that fixes the evidence source, numerical values, effect direction, and permissible language strength for each claim before a large language model (LLM) generates connective prose. This approach addresses failures where LLMs drift numbers or invert effect directions in scientific reports. Experiments on fMRI functional‑connectivity and randomized controlled trial reporting show that claim‑locked reporting improves reproducibility by 37.4 and 20.5 points over a deterministic hybrid template, while also reducing token usage and generation latency.

By Xiao Fan, Jingyuan Li, Hongbin Guo, Yubo Han, Yi Zhang
arXiv AI
4d ago

Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis

The paper introduces the Active Provenance Gate (APG), a post‑debate verification layer for multi‑agent debate synthesis that audits debate logs, applies self‑correction, and blocks unsupported claims before publication. Empirical studies show that APG more than doubles provenance fidelity in crisis simulations and that users prefer explicit failure reports over fabricated consensus. The work shifts data origin tracing from passive logging to active conditional blocking, addressing safety gaps in large‑language‑model‑based debate systems.

By Jakub Mas{\l}owski, Jaros{\l}aw A. Chudziak
arXiv Computation and Language
Aug 21

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

arXiv:2608. 20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence.

By Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave, Xiao Gu, David A. Clifton
arXiv AI
Jun 2

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.

By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu