Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway.
The paper discusses how large enterprises can adopt the harness paradigm to overcome limitations of traditional coding approaches. It reviews recent findings that harnesses outperform complex architectures at the task level, that harness choice drives benchmark variance more than model choice, and that governance is the main barrier to enterprise adoption. The authors propose a unified harness architecture that keeps code identical across deployments, simplifying review and audit processes.
By George Juraj Salapa
The paper discusses how enterprises increasingly deploy AI coding agent harnesses, often purchased from vendors like Anthropic or OpenAI, and how these harnesses dictate model choice, prompt handling, and cost. It introduces a fast, customizable routing system that classifies prompts and strategically routes them to minimize expensive model usage, achieving 14–21% cost savings in a simulated 10,000-seat enterprise. The study also evaluates risks across twenty harnesses, highlights vendor dependence, and proposes an internal control plane for future harness ownership decisions.
By Arian Abbasi, Alan Aqrawi, Ted Kwartler
arXiv:2608. 08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the \emph{harness}---is typically treated as a fixed artifact after deployment.
By Tailin Zhou
The paper introduces HARNESSEVO, a method that decomposes a large language model’s harness into four independently evolvable components—role, task‑strategy, tool/format‑rules, and reflection/control. Experiments on ALFWorld show that overall success rates are similar to flat‑string evolution, but the reflection/control component alone accounts for most of the performance gains. The study also finds that evenly distributing optimization budget across all slots can be detrimental; concentrating resources on the high‑credit control slot recovers lost performance, while on WebShop all slots remain ineffective, suggesting task‑specific differences in harness value.
By Michael Nguyen, Wei Chen Tan, Nurul Aisyah Hassan, Arvind Raman, Li Hua Lim, Ahmad Faiz Razak
arXiv:2609.01437v1 Announce Type: cross
Abstract: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly...
By Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang