arXiv AI By Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan

Towards a Science of AI Agent Reliability

Read the original on arXiv AI →

arXiv:2602. 16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

Locating Hidden Failures Makes Long-Horizon Agents More Reliable

The paper introduces Traverse, a benchmark of 2,518 agent trajectories and 6,967 annotated mistakes across software engineering, computer use, and science tasks, revealing that failures often go unrecovered and can cause irreversible harm before a run is deemed successful. It shows that human judges struggle to detect the first mistake in most runs, while a 4‑billion‑parameter verifier called Scout can locate failures more effectively and improve task success when used to select among candidate runs. The study demonstrates that making failure detection inexpensive and reliable can enable long‑horizon agents to learn from their own mistakes and increase trustworthiness in autonomous AI.

By Salman Rahman, Yubin Kim, Mihir Parmar, A. Ali Heydari, Genglin Liu, Simon A. Lee, Weizhi Zhang, Arian Hosseini, Ahmed A. Metwally, Yuzhe Yang, Baharan Mirzasoleiman, Xin Liu, Pavel Izmailov, Saadia Gabriel, Mark Malhotra, Shwetak Patel, Daniel McDuff, Hamid Palangi
arXiv AI
Sep 7

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Harbor Adapters is a unified evaluation infrastructure that ports over 80 agentic benchmarks, enabling arbitrary agents to be tested across complex environments. The authors performed a large‑scale evaluation of 8 models on 54 benchmarks, using Terminus‑2 and three native harnesses, revealing detailed agent capabilities and failure modes. They also created Harbor‑Index, a curated set of 82 challenging tasks from 29 benchmarks, designed to be affordable yet comprehensive, with the best model achieving a 28.0% pass rate.

By Lin Shi (Audrey), Haowei Lin (Audrey), Zixuan Zhu (Audrey), Xiaoyue Zhou (Audrey), Xiang Li (Audrey), Xiangning Lin (Audrey), Yaxuan Deng (Audrey), Han Xu (Audrey), Yuangang Li (Audrey), Shanda Li (Audrey), Zizhao Chen (Audrey), Hanwen Xing (Audrey), Harsh Raj (Audrey), Bo Chen (Audrey), Quan Shi (Audrey), Steven Dillmann (Audrey), Yipeng Gao (Audrey), Puneesh Khanna (Audrey), Ruofan Lu (Audrey), Chao Beyond Zhou (Audrey), Michael Yang (Audrey), Robert Zhang (Audrey), Siyuan Chai (Audrey), Jiayu Chang (Audrey), Yizhao Chen (Audrey), Xiaokun Chen (Audrey), Yiwei Dai (Audrey), Wenting Yang (Audrey), Hange Liu (Audrey), Minghao Liu (Audrey), Zihan Wang (Audrey), Adnan El Assadi (Audrey), Benedikt Stroebl (Audrey), E. Kelly Buchanan (Audrey), Han Meng (Audrey), Junwei He (Audrey), Longxuan Yu (Audrey), Radin Shayanfar (Audrey), Yukyung Lee (Audrey), Zhikang Dong (Audrey), Allen G Hart (Audrey), Anjiang Wei (Audrey), Anurag Kashyap (Audrey), Arpandeep Khatua (Audrey), Audrey Jixin Zheng (Audrey), Chengrui Ma (Audrey), David Heineman (Audrey), Dubing Chen (Audrey), Hai-Anh Trinh (Audrey), Haishuo Fang (Audrey), Hefan Zhang (Audrey), Hui Shen (Audrey), Issa Sugiura (Audrey), Jiankai Sun (Audrey), Jiechao Gao (Audrey), Junhong Lin (Audrey), Junnan Li (Audrey), Kai Yang (Audrey), Lei Hsiung (Audrey), Maoyu Wang (Audrey), Mengze Tang (Audrey), Nabil Omi (Audrey), Negin Raoof (Audrey), Nicholas Edwards (Audrey), Octavia Guo (Audrey), Orfeas Menis Mastromichalakis (Audrey), Pengliang Ji (Audrey), Przemys{\l}aw Hejman (Audrey), Qi Qi (Audrey), Qunshu Lin (Audrey), Richard Zhuang (Audrey), Rui Yang (Audrey), Ruichen Zheng (Audrey), Ryan Marten (Audrey), Shaghayegh Fazliani (Audrey), Shizheng Hou (Audrey), Sicong Jiang (Audrey), Sijie Li (Audrey), Song Bian (Audrey), Terry Yue Zhuo (Audrey), Tianqing Wu (Audrey), Tom Tang (Audrey), Wanjia Zhao (Audrey), Weihao Xuan (Audrey), Wenhua Liang (Audrey), Xian Liu (Audrey), Xin Lan (Audrey), Xuan Zhang (Audrey), Xuandong Zhao (Audrey), Yanchuan Tang (Audrey), Yifan Jiang (Audrey), Yijiang Li (Audrey), Yitong Guan (Audrey), Yizhi Li (Audrey), Yonghui Liu (Audrey), Yuheng Tang (Audrey), Yujun (Audrey), Mao, Yunfei Zhao, Yuxin Wang, Yuxuan Tang, Zhenheng Tang, Zhifei Li, Ziruo Wang, Ziyu She, Kaiyuan Liu, Iheb Chaabane, Yuxin Tang, Xiangyi Li, Andy Konwinski, Boxuan Li, Leon Liangyu Chen, Alex Dimakis, Nicholas Carlini, Soroush Vosoughi, Di He, Etash Guha, Benjamin Feuer, Mike Merrill, Ludwig Schmidt, Alex Shaw
arXiv AI
2d ago

Incident-Arena: Getting agents to the last nine of reliability

Incident‑Arena is a new benchmark for AI coding agents focused on production incident response, featuring 20 tasks derived from real‑world open‑source software. Each task deploys a production application on an ephemerally created Kubernetes cluster, injects faults at various layers, and applies a sustained load profile. The benchmark introduces functional verifiers that maintain system‑level metrics while ensuring safe repairs, and shows that current frontier models achieve below 64.3% across the tasks, highlighting challenges in diagnosis, repair, and regression safety.

By Andre Fu, Malik Drabla, Leon Liu, Meji Abidoye, Marek Suppa, Lata Mishra, Adnan El Assadi, Yiyuan Li
arXiv AI
Sep 3

READY or Not: Reliable Enterprise Agent Deployment

READY or Not: Reliable Enterprise Agent Deployment introduces a framework for qualifying AI agents for enterprise workflows. It measures reliability and operating cost under various oversight policies, selects the minimum‑cost policy that meets a specified reliability target, and statistically qualifies it on held‑out cases. In a clinical audit study, READY revealed that two agents with nearly identical autonomous accuracy required markedly different levels of human review to achieve the same reliability target.

By Veronica Chatrath (Christy), Bryan Zhu (Christy), Jingxuan Fan (Christy), George Pu (Christy), Soham Dinesh Tiwari (Christy), Soham Dan (Christy), Ryan Young (Christy), Yuan (Christy), Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan Xue