Provenance Before Prose: Claim-Locked Reporting for Statistical Text Generation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces claim‑locked reporting, a protocol that fixes the evidence source, numerical values, effect direction, and permissible language strength for each claim before a large language model (LLM) generates connective prose. This approach addresses failures where LLMs drift numbers or invert effect directions in scientific reports. Experiments on fMRI functional‑connectivity and randomized controlled trial reporting show that claim‑locked reporting improves reproducibility by 37.4 and 20.5 points over a deterministic hybrid template, while also reducing token usage and generation latency.
arXiv:2609.00654v1 Announce Type: new Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scien...
The paper introduces the Active Provenance Gate (APG), a post‑debate verification layer for multi‑agent debate synthesis that audits debate logs, applies self‑correction, and blocks unsupported claims before publication. Empirical studies show that APG more than doubles provenance fidelity in crisis simulations and that users prefer explicit failure reports over fabricated consensus. The work shifts data origin tracing from passive logging to active conditional blocking, addressing safety gaps in large‑language‑model‑based debate systems.
arXiv:2608. 20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence.
arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.