Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take, which failures are repaired, which states are verified, and which evidence is logged. We show that this harness can change the agent's multi-step beliefs even when the task, environment, and base LLM are fixed.
arXiv:2606. 10241v1 Announce Type: new Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlogged, diagnoses cannot be replayed, and promote-or-discard decisions land in a side database rather than the agent's own history.
By Yohei Nakajima
arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.
By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
arXiv:2607. 26598v1 Announce Type: cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactions.
By Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu, Shanwen Tan, Bing Zhao, Boyu Yang, Zhijie Xu, Ming Kong, Hu Wei, Jie Liu, Qiang Zhu
The Agent Error Dataset (AED) presents 50,228 error–diagnosis pairs collected from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text‑based agent systems. A five‑stage Agentic Error‑to‑Training (AET) pipeline generates diagnoses and proposed corrections, verifies them against recorded evidence, and creates separate training views for diagnosis and actor recovery. Experiments show that first‑proposal corrections improve verifier pass rates from 18.4% to 51.1%, and fine‑tuning with full‑diagnosis data raises Qwen3‑8B’s exact‑step agreement from 47.2% to 63.6% on a holdout set.
By Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
arXiv:2609.16302v1 Announce Type: cross
Abstract: When a coding agent returns to existing software, it inherits evidence from earlier engineering work: tests, type checks, proofs, static analyses, an...
By Anjan Goswami
arXiv:2610.00948v1 Announce Type: cross
Abstract: The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termi...
By Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai
Mid‑Harness proposes a test‑time compute strategy that samples and verifies candidate actions before execution, keeping the underlying generator and harness unchanged. Experiments show that with a strong verifier, sampling more actions significantly boosts success rates—e.g., a GPT‑5.6 verifier raises Pass@1 from 50.00 % to 68.03 % on TerminalBench‑Lite using eight samples. The approach also improves performance across various models, benchmarks, and harnesses, demonstrating that action scaling is a promising target for enhancing terminal agent reliability.
By Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee
arXiv:2603. 14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions.
By Shengda Fan, Xuyan Ye, Yupeng Huo, Zhi-Yuan Chen, Yiju Guo, Shenzhi Yang, Wenkai Yang, Shuqi Ye, Jingwen Chen, Haotian Chen, Xin Cong, Yankai Lin
LEGO-RL is a framework that connects native coding-agent harnesses with scalable policy‑gradient training without altering the harnesses’ internal flow. It achieves faithful optimization through in‑process LLM proxying, reliable execution via sandbox orchestration, and observable training with automated validation and a Live UI. Experiments show LEGO‑RL improves the Qwen3.5‑35B‑A3B model’s performance on three native harnesses while preserving high rollout‑training probability correlation.
By Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
arXiv:2606. 29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit.
By Aojie Yuan, Yi Nian, Haiyue Zhang, Zijian Su, Yue Zhao
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion.