arXiv:2607. 22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified.
By Naman Vats, Oleg Golev
arXiv:2606. 12344v1 Announce Type: new Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.
By Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao Xu, Guohao Dai, Lixue Xia, Yunchao Wei, Yunhe Wang, Yu Wang
arXiv:2609.11987v1 Announce Type: cross
Abstract: An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous softwar...
By Mohsen Arjmandi
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring. We introduce Claw-SWE-Bench, a multilingual SWE-bench-style benchmark and adapter protocol that makes heterogeneous agent harnesses, or claws, comparable under fair settings including a fixed prompt, runtime budget, workspace contract, patch extraction procedure, and evaluator.
Mid‑Harness proposes a test‑time compute strategy that samples and verifies candidate actions before execution, keeping the underlying generator and harness unchanged. Experiments show that with a strong verifier, sampling more actions significantly boosts success rates—e.g., a GPT‑5.6 verifier raises Pass@1 from 50.00 % to 68.03 % on TerminalBench‑Lite using eight samples. The approach also improves performance across various models, benchmarks, and harnesses, demonstrating that action scaling is a promising target for enhancing terminal agent reliability.
By Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee
Mingbird is a local‑first agent harness designed for small open‑weight language models (2–9 B) that run on ordinary laptops. It introduces ten mechanisms—such as a byte‑level net‑zero prefill budget, a finish gate that re‑reads the task before accepting completion, and signature‑level loop detection—to address common failure modes that arise from the harness rather than the model itself. In controlled experiments on the LRAB benchmark and the $ au^2$‑bench, Mingbird achieves higher overall scores (0.886 and 0.856 respectively) compared to other harnesses, and its ablation studies show that each mechanism contributes measurable performance gains.
By Hao Wang, Ting Huang
The paper reports that applying reinforcement learning (RL) to the Kimi K2.7 Code model on 1,700 agentic coding tasks improves its performance on six external benchmarks. After a single epoch of GSPO training on a rank‑32 LoRA adapter, pass‑@1 scores increased across all benchmarks, with significant gains even on data released after training. The trained model also reduces agent steps and avoids common failure modes such as dropping requirements or breaking existing behavior.
By Sushant Mehta, Logan Ritchie, Edwin Chen
arXiv:2607. 12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents.
By Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao
arXiv:2608. 15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps.
By Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang
arXiv:2606. 17454v1 Announce Type: new Abstract: AI agent performance is not just a modeling problem, it is fundamentally a systems problem.
By Gaurav Gupta, Vatshank Chaturvedi, Jun Huan, Anoop Deoras
arXiv:2609.24974v1 Announce Type: cross
Abstract: Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain...
By Haoran Ye, Yuxing Lu, Haonan Dong, Zhaochen Su, Guojie Song
MoMHa is a system that optimizes large language model harnesses across three objectives—accuracy, behavioural safety, and token cost—using a single‑phase joint‑reward proposer. It outperforms alternative strategies on seventeen domains, including synthetic suites and real‑world benchmarks, achieving higher joint scores and better safety while reducing token usage. The approach demonstrates that multi‑objective harness design can transfer effectively to unseen models and tasks.
By Subhojyoti Mukherjee, Md Mehrab Tanjim