arXiv:2607. 22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified.
By Naman Vats, Oleg Golev
arXiv:2607. 10569v1 Announce Type: cross Abstract: Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters.
By Hong Yang, Qi Yu, Travis Desell
arXiv:2607. 03691v1 Announce Type: cross Abstract: Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops.
By Oussama Ben Sghaier, Hao Li, Bram Adams, Ahmed E. Hassan
arXiv:2607. 28545v2 Announce Type: replace-cross Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.
By Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi
arXiv:2608.28021v2 Announce Type: replace-cross
Abstract: Large language models increasingly author Infrastructure-as-Code (IaC), where one insecure default is provisioned straight into production. P...
By Animesh Shaw
Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses.
Mingbird is a local‑first agent harness designed for small open‑weight language models (2–9 B) that run on ordinary laptops. It introduces ten mechanisms—such as a byte‑level net‑zero prefill budget, a finish gate that re‑reads the task before accepting completion, and signature‑level loop detection—to address common failure modes that arise from the harness rather than the model itself. In controlled experiments on the LRAB benchmark and the $ au^2$‑bench, Mingbird achieves higher overall scores (0.886 and 0.856 respectively) compared to other harnesses, and its ablation studies show that each mechanism contributes measurable performance gains.
By Hao Wang, Ting Huang
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring. We introduce Claw-SWE-Bench, a multilingual SWE-bench-style benchmark and adapter protocol that makes heterogeneous agent harnesses, or claws, comparable under fair settings including a fixed prompt, runtime budget, workspace contract, patch extraction procedure, and evaluator.
AgentRoom introduces a real‑time collaborative editing protocol that enables concurrent coding by multiple large language model agents within a CRDT‑backed shared workspace. By providing file‑level claim, status, and broadcast tools, it allows agents to coordinate directly rather than relying on serial phase handoffs or independent sampling. Experiments with five frontier coding‑CLI models show that AgentRoom reduces task abandonment and run‑to‑run variation compared to solo or parallel‑merge approaches, highlighting the importance of coordination over mere parallelism.
By Seonglae Cho, Donghyun Lee
arXiv:2608.28795v1 Announce Type: cross
Abstract: Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screensho...
By Achint Mehta
arXiv:2609.11987v1 Announce Type: cross
Abstract: An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous softwar...
By Mohsen Arjmandi
arXiv:2606. 12344v1 Announce Type: new Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.
By Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao Xu, Guohao Dai, Lixue Xia, Yunchao Wei, Yunhe Wang, Yu Wang