arXiv AI

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent

TimeEvo is a new method for time‑series agents that autonomously evolves its tool library based on failures observed during runtime. By clustering diagnosed failures into capability gaps, planning measurements, synthesizing evidence‑only tools, and admitting candidates through a paired gate, the system starts from an empty library and improves accuracy across ten QA tasks and three backbones. Experiments show that even a library built on a cheap model benefits stronger models when installed.

arXiv Machine Learning
Sep 17

Locating Hidden Failures Makes Long-Horizon Agents More Reliable

The paper introduces Traverse, a benchmark of 2,518 agent trajectories and 6,967 annotated mistakes across software engineering, computer use, and science tasks, revealing that failures often go unrecovered and can cause irreversible harm before a run is deemed successful. It shows that human judges struggle to detect the first mistake in most runs, while a 4‑billion‑parameter verifier called Scout can locate failures more effectively and improve task success when used to select among candidate runs. The study demonstrates that making failure detection inexpensive and reliable can enable long‑horizon agents to learn from their own mistakes and increase trustworthiness in autonomous AI.

By Salman Rahman, Yubin Kim, Mihir Parmar, A. Ali Heydari, Genglin Liu, Simon A. Lee, Weizhi Zhang, Arian Hosseini, Ahmed A. Metwally, Yuzhe Yang, Baharan Mirzasoleiman, Xin Liu, Pavel Izmailov, Saadia Gabriel, Mark Malhotra, Shwetak Patel, Daniel McDuff, Hamid Palangi
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
2d ago

Revision-Aware Independent Agent Graphs for Dynamic Reasoning

The paper introduces Revision‑Aware Independent Agent Graphs (RIAG) to address dynamic task routing, where an event stream continually revises task bindings and a system must select the correct document version at query time. By repurposing six benchmarks into over 31,000 dynamic episodes, the authors demonstrate that RIAG balances recomputation and reuse, achieving 54.24 % joint routing‑and‑answer accuracy with only 0.62 calls per query—substantially better than the strongest baseline. The study highlights the trade‑off between stale conclusions and wasted work in dynamic reasoning settings.

By Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang
arXiv AI
Sep 3

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

ToolGate is an executable acceptance pipeline designed to streamline the creation of scientific benchmarks that rely on specialist software. It evaluates each model-generated item through three gates: (1) an executable solution script must reproduce the proposed answer, (2) a randomized no‑tool screen rejects items solvable without the software, and (3) a tool‑using agent must solve the item within a time limit. In a FEniCSx instantiation, 500 generation attempts produced 128 unique, verified benchmark items after successive filtering.

By Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi
arXiv AI
Sep 17

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

The paper argues that agentic systems waste time and memory by guessing how long tool calls will take, rather than using explicit progress signals from the tools themselves. It demonstrates that tools can report their remaining work or imminent completion, and that incorporating this feedback into serving systems dramatically improves cache decisions and reduces token latency. The authors show that this approach outperforms existing predictors and works robustly across different environments.

By Yipeng Liu, Yingqiang Zhang, Feifei Li, Huanchen Zhang
arXiv AI
Aug 11

$A^2E$ : An End-to-End Agent Auditing Engine

arXiv:2608. 07346v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou