arXiv AI

Verifiable abstention makes AI leak diagnosis accountable in water distribution networks

The paper proposes a verifiable abstention framework for AI-driven leak localization in water distribution networks, where a physics-based executor agent tests leak hypotheses against a digital twin and an independent supervisor with an LLM auditor certifies actions or abstains. In noisy field conditions, the system achieves 96% decision precision on acted events, correctly identifies all leaks in a benchmark, and demonstrates practical deployment with a 194-event audit record. The approach offers a defensible, accountable method for autonomous water‑infrastructure operation.

arXiv AI
2d ago

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness is a method that enhances verification for large language model agents tackling long‑horizon tasks without needing reference answers at test time. It transforms the base LLM into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills, using disagreement resolution and consensus challenge to evaluate competing claims. Across five benchmarks and two frontier models, VeriHarness outperforms baselines, achieving significant performance gains and demonstrating self‑improvement of verification skills from failure feedback.

By Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee
arXiv Machine Learning
Jul 28

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.

By Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele
arXiv AI
6d ago

Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI

SustainAI is a water‑aware, closed‑loop framework that embeds environmental accountability into AI deployment. It combines real‑time water metering, a hallucination‑aware penalty model, and a water‑aware routing algorithm that considers regional water stress. In tests with small language models, water footprints varied 11‑fold across data centers, and 1,335 inference runs consumed about 399 mL of water but yielded only 240 correct outputs, highlighting the resource cost of inaccurate responses.

By Farnaz Farid, Tashfia Towkee, Sania Nasreen, Sami bin Azad
arXiv AI
Sep 11

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

The paper introduces NetArtifactBench, a benchmark designed to evaluate whether AI agents can detect and repair inconsistencies in network experiment records while preserving supported claims. It tests 23 agent configurations on 52 instances with injected inconsistencies, finding an average pass rate of 65.3 % but no runtime exceeding 30 % for complex repairs that require recovering implicit relations and propagating changes across artifacts. The results highlight a clear distinction between local corrections and full record-level repair, leading the authors to argue that artifact integrity should be a primary design and evaluation criterion for AI agents in network systems.

By Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu