arXiv:2609.24974v1 Announce Type: cross
Abstract: Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain...
By Haoran Ye, Yuxing Lu, Haonan Dong, Zhaochen Su, Guojie Song
arXiv:2609.13739v1 Announce Type: cross
Abstract: Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory...
By Hongliang Wei (Harbin Institute of Technology, Alibaba Cloud), Xiaobing Tu (Alibaba Cloud), Yinggui Wang (Alibaba Cloud), Zhengxi Liu (Alibaba Cloud), Rongkun Xue (Alibaba Cloud), Jinkui Ren (Alibaba Cloud), Xiantao Zhang (Alibaba Cloud), Debin Zhao (Harbin Institute of Technology), Xiaopeng Fan (Harbin Institute of Technology)
arXiv:2609.06107v1 Announce Type: new
Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which do...
By Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang
The paper introduces CHART, a curriculum that rotates harnesses during training to teach search agents parallel search strategies robustly across different harness configurations. Unlike static harness augmentation, CHART gradually consolidates behavior by graduating learned harnesses and replacing them, maintaining a reward gap that drives learning. Experiments show CHART enables agents to parallelize on 89% of held‑out harnesses, improves performance on a new QA task by 5.6pp, and benefits more from meta‑harness search than baselines.
By Xinlu Zhang, Ying-Chun Lin, Zhihan Zhang, Besnik Fetahu, Xi Chen
arXiv:2609.11987v1 Announce Type: cross
Abstract: An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous softwar...
By Mohsen Arjmandi
Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability. We introduce LoRA Scaffolded Policy Optimization (LSPO), a sampling-time mechanism that recovers this lost gradient.