arXiv:2608.18836v2 Announce Type: replace
Abstract: Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation. Here, we qua...
By Tianwei Mu, Yue Wang, Mingzhe Yuan, Manhong Huang, Wenhong Wang, Xuerui Yin, Qing Luo, Min Xiao, Hui Yang, Jun Li, Dan Xue
arXiv:2608.22160v1 Announce Type: new
Abstract: Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at p...
By Zhixu Du, Yiran Chen
arXiv:2609.15319v1 Announce Type: cross
Abstract: Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced a...
By Luis M. S\'anchez
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
By Zeyu Zhang, Bradly C. Stadie
arXiv:2606. 19356v1 Announce Type: cross Abstract: When multi-agent LLM systems produce bad answers, not all failures are equal: some answers are grounded in the right material but incomplete, while others are simply ungrounded and should be stopped.
By Anantha Sharma
VeriHarness is a method that enhances verification for large language model agents tackling long‑horizon tasks without needing reference answers at test time. It transforms the base LLM into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills, using disagreement resolution and consensus challenge to evaluate competing claims. Across five benchmarks and two frontier models, VeriHarness outperforms baselines, achieving significant performance gains and demonstrating self‑improvement of verification skills from failure feedback.
By Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee
arXiv:2608. 02774v1 Announce Type: cross Abstract: AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal sensitive details about the model, workload, or hardware.
By Sleem Abdelghafar, Gabriel Kulp
arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.
By Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele
arXiv:2608. 15657v1 Announce Type: new Abstract: The rapid development and widespread application of artificial intelligence (AI) have sparked intense discussions on how to deploy responsible AI systems in a manner aligned with human values and ethical standards.
By Chong Chen, Yulu Zhang, Qingxi Guo, Yihan Liu
arXiv:2609.23360v1 Announce Type: cross
Abstract: Engineering predictions require physical mechanisms to be translated consistently into equations, discretization, code, and validation, yet errors ca...
By Jie Shi, Yimin Lu, Zhongkun Ouyang
SustainAI is a water‑aware, closed‑loop framework that embeds environmental accountability into AI deployment. It combines real‑time water metering, a hallucination‑aware penalty model, and a water‑aware routing algorithm that considers regional water stress. In tests with small language models, water footprints varied 11‑fold across data centers, and 1,335 inference runs consumed about 399 mL of water but yielded only 240 correct outputs, highlighting the resource cost of inaccurate responses.
By Farnaz Farid, Tashfia Towkee, Sania Nasreen, Sami bin Azad
The paper introduces NetArtifactBench, a benchmark designed to evaluate whether AI agents can detect and repair inconsistencies in network experiment records while preserving supported claims. It tests 23 agent configurations on 52 instances with injected inconsistencies, finding an average pass rate of 65.3 % but no runtime exceeding 30 % for complex repairs that require recovering implicit relations and propagating changes across artifacts. The results highlight a clear distinction between local corrections and full record-level repair, leading the authors to argue that artifact integrity should be a primary design and evaluation criterion for AI agents in network systems.
By Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu