The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
By Qiankai Xu
DAGent introduces an Evaluate‑then‑Grow planning approach for deep research agents, building directed acyclic graphs incrementally based on confidence and uncertainty from completed tasks. The framework includes a hierarchical context layer for efficient query handling and a structural reinforcement learning component, DAGRPO, that rewards topology‑conditioned execution. Experiments on BrowseComp‑Plus, GAIA, and xbench‑DeepSearch show DAGent outperforming strong baselines across multiple backbones and scaling to large language models.
By Hanwen Liu, Yuanfu Sun, Qiaoyu Tan
Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.
By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv:2609.39967v1 Announce Type: cross
Abstract: Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few...
By Yuliana Shakhvalieva, Dmitrii Kharchev, Viacheslav Bezrukov, Inessa Fedorova, Dmitry Bocharov, Ivan Oseledets, Valerii Ternovskii
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
By Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, Roi Reichart
arXiv:2608. 09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop.
By Hui Xue, Fan Yang