arXiv AI

Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

The paper introduces Provenance Density, an interface that visualizes the density of verified claims within a text to counter the Fluency Trap—where users mistake fluent AI-generated hallucinations for truth. In a study with 81 participants, the interface significantly improved users’ ability to distinguish true from fabricated content, while no signal led to no discernment. A technical audit of 200 samples revealed that retrieval density alone is insufficient, and that the Consistency Veto provides most of the discriminative power for dynamic queries.

arXiv Computation and Language
Aug 27

Provenance Before Prose: Claim-Locked Reporting

The paper introduces claim‑locked reporting, a protocol that fixes the evidence source, numerical values, effect direction, and permissible language strength for each claim before a large language model (LLM) generates connective prose. This approach addresses failures where LLMs drift numbers or invert effect directions in scientific reports. Experiments on fMRI functional‑connectivity and randomized controlled trial reporting show that claim‑locked reporting improves reproducibility by 37.4 and 20.5 points over a deterministic hybrid template, while also reducing token usage and generation latency.

By Xiao Fan, Jingyuan Li, Hongbin Guo, Yubo Han, Yi Zhang
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang