EngramBench is a new benchmark designed to evaluate skill evolution in autonomous agents by focusing on genuine capability abstraction rather than solution copying. It includes 30 learning tasks and 13 unseen transfer tasks that require agents to manage complex, multi-hour development cycles with LLM‑simulated users. The study shows that while static skill banks cannot eliminate the need for precise code implementation, they effectively reduce redundant context and cut overall coding time by more than 55%.
By Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao
arXiv:2606. 10087v1 Announce Type: cross Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats.
By Ankit Gupta, Aditya Prasad, Rameswar Panda
The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
By Qiankai Xu
arXiv:2609.08435v2 Announce Type: new
Abstract: In persistent interactions, long contexts may encode an evolving process rather than a fixed record: later events can revise or revoke earlier informat...
By Ziliang Zhao, Zenan Xu, Shuting Wang, Zhao Wang, Bowen Cao, Minda Hu, Lincheng Li, Pluto Zhou, Zhicheng Dou
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
By Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, Roi Reichart
LongWoF-Bench is a new benchmark of 778 machine‑verifiable long‑workflow tasks spanning code generation, agent‑environment synthesis, mathematical reasoning, and rule following. The study shows that EvoMap Genes—structured representations of verifier‑confirmed execution trajectories—outperform the Skill baseline by 8.7–15.5 percentage points across seven models, and for Claude Opus they enable 39 additional task completions while cutting token consumption by 9.9%. The results demonstrate that verified execution experience can be externalized and reused, improving long‑workflow completion without repeatedly discovering new strategies.
By Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
arXiv:2607. 27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation.
By Yihao Chen, Shi Chang, Khaled Chawa, Feng Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
arXiv:2609.14857v1 Announce Type: new
Abstract: Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve executio...
By Siwei Wu, Jincheng Ren, Yizhi Li, Haau-Sing Li, Chengran Yang, Yuxuan Zhang, Weicheng Gu, Jian Yang, Riza Batista-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zhou, Bryan Dai, Chenghua Lin
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end veri...
arXiv:2604. 18543v4 Announce Type: replace Abstract: Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale.
By Xirui Li, Ming Li, Ion Stoica, Cho-Jui Hsieh, Tianyi Zhou
arXiv:2601. 03808v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handled via heuristic design or brute-force approaches.
By Usha Shrestha, Dmitry Ignatov, Radu Timofte
arXiv:2607. 25675v1 Announce Type: new Abstract: Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box.
By Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao