arXiv AI By Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Read the original on arXiv AI →

arXiv:2607. 14004v1 Announce Type: new Abstract: Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler introduces automated curriculum learning for harness optimization, treating the evolving training curriculum as a non‑stationary bandit problem. It identifies reusable failure patterns, estimates learning progress for each, and balances revisiting known weaknesses with exploring new scenarios, allowing the curriculum to co‑evolve with the harness. Experiments on GAIA2 and Terminal‑Bench 2.0 show consistent improvements in harness performance, with Pass@1 gains of 4.4 and 7.5 percentage points over fixed‑order baselines.

By Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor R\"uhle
arXiv Computation and Language
Sep 7

EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

EVOHARNESSBENCH is a new benchmark that tests how LLM-based agents handle changes in their tool, skill, and agent harnesses over time. It includes 17 deterministic harness streams with 802 tasks, 520 tools, 42 skills, and 62 agents, and evaluates agents in two settings: deployment evaluation and self‑evolving adaptation evaluation. The study finds that harness expansion can cause forgetting, adaptation gains are inconsistent, and preserving old competence does not always aid new capability adaptation, highlighting harness evolution as a distinct challenge for agent development.

By Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
arXiv AI
3d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu