arXiv AI

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

arXiv:2606. 01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal.

arXiv AI
2d ago

EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

EngramBench is a new benchmark designed to evaluate skill evolution in autonomous agents by focusing on genuine capability abstraction rather than solution copying. It includes 30 learning tasks and 13 unseen transfer tasks that require agents to manage complex, multi-hour development cycles with LLM‑simulated users. The study shows that while static skill banks cannot eliminate the need for precise code implementation, they effectively reduce redundant context and cut overall coding time by more than 55%.

By Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao
arXiv AI
2d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu
arXiv Computation and Language
Aug 25

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

LongWoF-Bench is a new benchmark of 778 machine‑verifiable long‑workflow tasks spanning code generation, agent‑environment synthesis, mathematical reasoning, and rule following. The study shows that EvoMap Genes—structured representations of verifier‑confirmed execution trajectories—outperform the Skill baseline by 8.7–15.5 percentage points across seven models, and for Claude Opus they enable 39 additional task completions while cutting token consumption by 9.9%. The results demonstrate that verified execution experience can be externalized and reused, improving long‑workflow completion without repeatedly discovering new strategies.

By Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
arXiv Computation and Language
Sep 15

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement

arXiv:2609.14857v1 Announce Type: new Abstract: Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve executio...

By Siwei Wu, Jincheng Ren, Yizhi Li, Haau-Sing Li, Chengran Yang, Yuxuan Zhang, Weicheng Gu, Jian Yang, Riza Batista-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zhou, Bryan Dai, Chenghua Lin
arXiv AI
Jul 29

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

arXiv:2607. 25675v1 Announce Type: new Abstract: Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box.

By Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao