arXiv Computation and Language

Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

The paper introduces HARNESSEVO, a method that decomposes a large language model’s harness into four independently evolvable components—role, task‑strategy, tool/format‑rules, and reflection/control. Experiments on ALFWorld show that overall success rates are similar to flat‑string evolution, but the reflection/control component alone accounts for most of the performance gains. The study also finds that evenly distributing optimization budget across all slots can be detrimental; concentrating resources on the high‑credit control slot recovers lost performance, while on WebShop all slots remain ineffective, suggesting task‑specific differences in harness value.

arXiv AI
Sep 10

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

arXiv:2609.09134v1 Announce Type: new Abstract: Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic...

By Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur
arXiv AI
Jul 9

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

arXiv:2607. 06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value.

By Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh
arXiv Machine Learning
Sep 22

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

The paper introduces Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), a method that applies regularization principles to the iterative editing of an LLM agent’s harness—prompts, control flow, tooling, memory, and context management. RRSI limits the number of edits per candidate, encourages novel trajectories, and uses a critic and pruner to filter out benchmark‑specific or ineffective changes, thereby favoring reusable agent mechanisms. Experiments on eight benchmarks show RRSI improves performance by up to 14.1 points on the training split and 4.7 points on out‑of‑distribution tests, while reducing policy token usage by 30% compared to unregularized evolution.

By Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
arXiv AI
1d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu
arXiv AI
4d ago

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

MoMHa is a system that optimizes large language model harnesses across three objectives—accuracy, behavioural safety, and token cost—using a single‑phase joint‑reward proposer. It outperforms alternative strategies on seventeen domains, including synthetic suites and real‑world benchmarks, achieving higher joint scores and better safety while reducing token usage. The approach demonstrates that multi‑objective harness design can transfer effectively to unseen models and tasks.

By Subhojyoti Mukherjee, Md Mehrab Tanjim
arXiv Computation and Language
Sep 2

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

arXiv:2609.01437v1 Announce Type: cross Abstract: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly...

By Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang