arXiv:2609.24974v1 Announce Type: cross
Abstract: Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain...
By Haoran Ye, Yuxing Lu, Haonan Dong, Zhaochen Su, Guojie Song
The paper introduces WHALE, a method that alternates between updating a language model’s weights and searching for a better harness (the code that manages context and control flow). By iteratively fine‑tuning the model under the current harness and then optimizing the harness under the updated model, WHALE improves performance across search QA, math reasoning, and chess puzzles, outperforming weight‑only, harness‑only, and Fast‑Slow Training by 4.15–24.38 percentage points in mean@8 accuracy. The approach uses either fixed phase lengths or an adaptive patience rule to decide when to switch phases, and the authors provide code on GitHub.
By Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn, Kangwook Lee
EVOHARNESSBENCH is a new benchmark that tests how LLM-based agents handle changes in their tool, skill, and agent harnesses over time. It includes 17 deterministic harness streams with 802 tasks, 520 tools, 42 skills, and 62 agents, and evaluates agents in two settings: deployment evaluation and self‑evolving adaptation evaluation. The study finds that harness expansion can cause forgetting, adaptation gains are inconsistent, and preserving old competence does not always aid new capability adaptation, highlighting harness evolution as a distinct challenge for agent development.
By Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
arXiv:2609.04280v2 Announce Type: replace-cross
Abstract: Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what the...
By Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
arXiv:2609.13739v1 Announce Type: cross
Abstract: Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory...
By Hongliang Wei (Harbin Institute of Technology, Alibaba Cloud), Xiaobing Tu (Alibaba Cloud), Yinggui Wang (Alibaba Cloud), Zhengxi Liu (Alibaba Cloud), Rongkun Xue (Alibaba Cloud), Jinkui Ren (Alibaba Cloud), Xiantao Zhang (Alibaba Cloud), Debin Zhao (Harbin Institute of Technology), Xiaopeng Fan (Harbin Institute of Technology)
arXiv:2607. 26598v1 Announce Type: cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactions.
By Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu, Shanwen Tan, Bing Zhao, Boyu Yang, Zhijie Xu, Ming Kong, Hu Wei, Jie Liu, Qiang Zhu