arXiv AI

Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

arXiv:2607. 13285v1 Announce Type: new Abstract: The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution.

arXiv AI
Jun 15

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

arXiv:2606. 14249v1 Announce Type: new Abstract: AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts.

By Tingyang Chen, Shuo Lu, Kang Zhao, Weicheng Meng, Hanlin Teng, Tianhao Li, Chao Li, Xule Liu, Jian Liang, Zhizhong Zhang, Yuan Xie, Heng Qu, Kun Shao, Jian Luan
arXiv Computation and Language
Sep 15

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement

arXiv:2609.14857v1 Announce Type: new Abstract: Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve executio...

By Siwei Wu, Jincheng Ren, Yizhi Li, Haau-Sing Li, Chengran Yang, Yuxuan Zhang, Weicheng Gu, Jian Yang, Riza Batista-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zhou, Bryan Dai, Chenghua Lin
arXiv AI
3d ago

Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives

The paper introduces Harness Primitives—reusable agent harness mechanisms mined from failed task trajectories—and a framework called STITCH that selects and composes these primitives into task‑specific harnesses at test time. This approach avoids generating or debugging harness code for each task, achieving up to 12‑point gains in task success over fixed harness baselines and outperforming human‑designed harnesses like Codex CLI. STITCH also demonstrates minimal test‑time overhead (2.7%) and scales efficiently with the size of the primitive library.

By Peng Kuang, Haibo Jin, Dehao Wu, Feiyang Deng, Xiaopeng Yuan, Jerry Wang, Haohan Wang
arXiv AI
Sep 18

An Empirical Study of Harness Design for Coding Agents

The study investigates how individual components of a coding harness—planning, action space, and context management—affect autonomous coding agents’ performance. By fixing the execution loop and varying these components across 176 settings on SWE‑Bench Verified and Terminal‑Bench 2.1, the authors find that context management is most valuable when context windows are tight, staging rule‑based elision before LLM summarization yields the best efficiency, planning serves as an accuracy scaffold for weaker models and a cost saver for stronger ones, and predefined tools help models with limited bash skills while bash‑capable models benefit from a bash‑only interface. Trajectory‑level analysis shows that context management lengthens execution paths, planning alters where trajectories terminate, and the action space determines code granularity, offering a modular framework for future harness design.

By Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang
arXiv AI
3d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu
arXiv AI
Aug 28

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification

The paper introduces HarnessLens, a budget‑aware framework that evolves language‑model agent harnesses by jointly exploring task spaces and user‑configurable components. It derives candidate modifications from execution trajectories and selectively verifies them on behavior‑relevant tasks using an attributable‑evidence gate, thereby avoiding wasteful rollouts on unrelated behaviors. Experiments on three harnesses and four benchmarks show that HarnessLens improves held‑out performance by 7.6‑13.6% while using less evaluation budget than existing methods.

By Jinghan Xu, Yikai Zhang, Aili Chen, Weiyuan Li, Jiaqing Liang, Deqing Yang