arXiv AI By Tianwei Mu, Yue Wang, Mingzhe Yuan, Manhong Huang, Wenhong Wang, Xuerui Yin, Qing Luo, Min Xiao, Hui Yang, Jun Li, Dan Xue

Verifiable abstention makes AI leak diagnosis accountable in water distribution networks

Read the original on arXiv AI →

The paper proposes a verifiable abstention framework for AI-driven leak localization in water distribution networks, where a physics-based executor agent tests leak hypotheses against a digital twin and an independent supervisor with an LLM auditor certifies actions or abstains. In noisy field conditions, the system achieves 96% decision precision on acted events, correctly identifies all leaks in a benchmark, and demonstrates practical deployment with a 194-event audit record. The approach offers a defensible, accountable method for autonomous water‑infrastructure operation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness is a method that enhances verification for large language model agents tackling long‑horizon tasks without needing reference answers at test time. It transforms the base LLM into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills, using disagreement resolution and consensus challenge to evaluate competing claims. Across five benchmarks and two frontier models, VeriHarness outperforms baselines, achieving significant performance gains and demonstrating self‑improvement of verification skills from failure feedback.

By Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee