HarnessEvolve is a self‑evolving framework that improves agent harnesses—prompts, skills, tools, and execution logic—by learning from reference trajectories. It separates execution, evaluation, optimization, and gating into independent modules, addressing credit assignment failure, shortcut learning, and catastrophic forgetting. The approach uses reference trajectories to extract error signals, applies quality and performance gates to candidate updates, and validates updates on held‑out data, consistently outperforming state‑of‑the‑art baselines across diverse benchmarks.
By Wen Jiang, Mingmin Chu, Yimeng Tian, Qianxin Zhang, Haofei Yang, Rui Yang, Yang Liu, Tao Lv, Fangming Li
arXiv:2607. 05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.
By Xingze Gao, Chuanrui Hu, Hongda Chen, Pengfei Yao, Zhao Wang, Yi Bai, Zhengwei Wu, Yunyun Han, Xiaofeng Cong, Jie Gui, Yafeng Deng, Teng Li
Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer.
arXiv:2607. 21596v2 Announce Type: replace Abstract: Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution.
By Zeyu Ren, Ling Yue, Ran Li, Yishu Wang, Shengxiang Xu, Hanmo Liu, Shaowu Pan, Shimin Di
arXiv:2606. 17546v1 Announce Type: new Abstract: Self-evolving LLM-based agents improve mainly by changing their agent harness: the structured execution layer around a base model, including prompts, memory, tools, middleware, runtime state, and the model-tool interaction loop.
By Congjie Zheng, Chuanyi Xue, Bin Liang, Jun Yang, Changshui Zhang
The paper introduces a continuous evaluation framework that assesses both outcome-level and process-level aspects of evolving enterprise AI agent skills. It applies this framework to two variants of a Business Value Determination skill, running 240 trials across multiple models, harnesses, and specifications. The results show that while most trials pass final numerical checks, a large majority still exhibit process-level deviations, and dependency attribution reduces the number of failed checks per run. The framework also provides reusable regression tests and highlights specification sensitivity across configurations.
By Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin
The paper "It Takes Workflows to Evolve Better Workflows" introduces FloWright, a method that uses a hierarchical, structure‑aware reward system to allow one or more roles in a multi‑agent workflow to self‑evolve without extra models or data. It also proposes DataWright, an adaptive data hardening technique that transforms existing datasets into more challenging workflow‑level tasks. Experiments on document, slide, chart, code, math, and finance tasks show that small open models trained with FloWright can improve performance by up to +7.41%, with co‑evolving multiple roles yielding the largest gains.
"whyItMatters":"The work demonstrates that optimizing beyond the workflow generator—by enabling multiple agents to co‑evolve—can substantially enhance the effectiveness of multi‑agent workflows for complex real‑world tasks."
By Xuehang Guo, Haoyu Wang, Haifeng Chen, Yangyi Chen, Zhenhailong Wang, Qingyun Wang
The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
By Qiankai Xu
arXiv:2606. 15363v1 Announce Type: new Abstract: Self-improvement in AI agents has emerged as a key research frontier: systems that modify their own prompts, workflows, and decision rules based on accumulated operational experience.
By Ya-Chuan Chen, Tien-Jen Lai, Hsiang-Wei Hu
arXiv:2608. 04968v1 Announce Type: new Abstract: The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from failure.
By Jun Nie, Yonggang Zhang, Qianshu Cai, Yiu-ming Cheung, Xinmei Tian, Bo Han
arXiv:2609.24663v1 Announce Type: new
Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As t...
By Hongqiang Lin, Chao Liu, Xiaofan Bai, Xuan Jin, Yuhong Li, Nenggan Zheng, Xipeng Cao
arXiv:2607. 02469v1 Announce Type: cross Abstract: Software tests and code evolve together: a code change should be followed by new or updated tests that record the new software behavior.
By Jiale Amber Wang, Kaiyuan Wang, Pengyu Nie