The Agent Error Dataset (AED) presents 50,228 error–diagnosis pairs collected from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text‑based agent systems. A five‑stage Agentic Error‑to‑Training (AET) pipeline generates diagnoses and proposed corrections, verifies them against recorded evidence, and creates separate training views for diagnosis and actor recovery. Experiments show that first‑proposal corrections improve verifier pass rates from 18.4% to 51.1%, and fine‑tuning with full‑diagnosis data raises Qwen3‑8B’s exact‑step agreement from 47.2% to 63.6% on a holdout set.
By Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
arXiv:2607. 13884v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations.
By Wenjun Wang, Yuchen Fang, Fengrui Liu, Zibo Liang, Kai Zheng
The paper introduces FRESH, a Failure-aware Retrieval framework that uses Experience-Structured Heterogeneous graphs to transform past successes and failures into structured external memory for tool‑using agents. By explicitly modeling dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH enables frozen language models to reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on τ‑Bench and AppWorld with multiple open‑source models demonstrate that FRESH consistently improves task success and tool‑use reliability compared to no‑memory agents and other memory‑based baselines.
By Jiaxing Li, Lei Song, Rui Dong, Youyong Kong
arXiv:2606. 07412v1 Announce Type: cross Abstract: LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-quality SWE tasks.
By Chuan Xiao, Zhengbo Jiao, Shaobo Wang, Wei Wang, Bing Zhao, Hu Wei, Linfeng Zhang, Lin Qu
The paper introduces RUPA, a trajectory‑level uncertainty quantification framework for large language model agents. RUPA models an agent’s execution as a directed graph of reasoning states, tool interactions, and environment feedback, then propagates uncertainty across this graph to capture long‑range dependencies. Experiments on benchmarks such as τ‑2, Terminal‑Bench‑2, and GAIA show that RUPA outperforms existing methods, enabling earlier failure detection and more reliable agent execution.
By Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
arXiv:2606. 07127v1 Announce Type: new Abstract: Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed.
By Hikaru Shindo, Yu Deng, Teng Cao, Quentin Delfosse, Christopher Tauchmann, Jannis Bl\"uml, Gopika Sudhakaran, Kristian Kersting