arXiv AI

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

The paper introduces WHALE, a method that alternates between updating a language model’s weights and searching for a better harness (the code that manages context and control flow). By iteratively fine‑tuning the model under the current harness and then optimizing the harness under the updated model, WHALE improves performance across search QA, math reasoning, and chess puzzles, outperforming weight‑only, harness‑only, and Fast‑Slow Training by 4.15–24.38 percentage points in mean@8 accuracy. The approach uses either fixed phase lengths or an adaptive patience rule to decide when to switch phases, and the authors provide code on GitHub.

arXiv Machine Learning
Sep 22

CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents

The paper introduces CHART, a curriculum that rotates harnesses during training to teach search agents parallel search strategies robustly across different harness configurations. Unlike static harness augmentation, CHART gradually consolidates behavior by graduating learned harnesses and replacing them, maintaining a reward gap that drives learning. Experiments show CHART enables agents to parallelize on 89% of held‑out harnesses, improves performance on a new QA task by 5.6pp, and benefits more from meta‑harness search than baselines.

By Xinlu Zhang, Ying-Chun Lin, Zhihan Zhang, Besnik Fetahu, Xi Chen
arXiv AI
Sep 15

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning

arXiv:2609.13739v1 Announce Type: cross Abstract: Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory...

By Hongliang Wei (Harbin Institute of Technology, Alibaba Cloud), Xiaobing Tu (Alibaba Cloud), Yinggui Wang (Alibaba Cloud), Zhengxi Liu (Alibaba Cloud), Rongkun Xue (Alibaba Cloud), Jinkui Ren (Alibaba Cloud), Xiantao Zhang (Alibaba Cloud), Debin Zhao (Harbin Institute of Technology), Xiaopeng Fan (Harbin Institute of Technology)
arXiv Computation and Language
Aug 25

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

arXiv:2608.15763v2 Announce Type: replace Abstract: AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low lat...

By TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin, Yongdong Luo, Yibo Hu, Meiguang Jin, Junfeng Ma, Weihang Pan, Jiaxin Zhao, Zulong Chen
arXiv AI
3d ago

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

MoMHa is a system that optimizes large language model harnesses across three objectives—accuracy, behavioural safety, and token cost—using a single‑phase joint‑reward proposer. It outperforms alternative strategies on seventeen domains, including synthetic suites and real‑world benchmarks, achieving higher joint scores and better safety while reducing token usage. The approach demonstrates that multi‑objective harness design can transfer effectively to unseen models and tasks.

By Subhojyoti Mukherjee, Md Mehrab Tanjim
arXiv AI
Sep 7

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

The study investigates how multi‑harness reinforcement learning (RL) affects coding agents by comparing two grouping strategies—Within (one group per task‑harness pair) and Cross (harnesses pooled within a task)—using a Qwen3‑8B policy trained on frozen task‑harness records from Aider, OpenHands, Qwen Code, and SWE‑agent. Across 24,000 sealed evaluations, the choice of evaluation harness dramatically increases solve rates (from 2.14 % to 9.27 %), while the grouping rule has a negligible effect. Both grouping rules yield similar gains on the same source harness, and Cross‑harness credit does not improve portability beyond Within‑harness credit, suggesting that multi‑harness RL reports should specify grouping boundaries and test on unseen harnesses.

By Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen
arXiv Machine Learning
Aug 27

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

JIT‑Agent is a model that automatically generates task‑adaptive agent harnesses for any off‑the‑shelf LLM, replacing manual, task‑specific harness design. It learns to compose, repair, and evolve harnesses using a fixed four‑module protocol, and its use boosts performance on benchmarks such as DeepSearchQA and OdysseyBench, outperforming several mature agent runtimes. The approach demonstrates that harness intelligence can be trained, transferred, and compounded independently of model scaling.

By Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie, Zhaochen Yu, Zihang Liu, Zhongxiang Sun, Qiankun Li, Yue Liao, Heng Chang, Xiaobin Hu, Qibing Ren, Wangchunshu Zhou, Shuicheng Yan