arXiv AI

Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

The paper discusses how enterprises increasingly deploy AI coding agent harnesses, often purchased from vendors like Anthropic or OpenAI, and how these harnesses dictate model choice, prompt handling, and cost. It introduces a fast, customizable routing system that classifies prompts and strategically routes them to minimize expensive model usage, achieving 14–21% cost savings in a simulated 10,000-seat enterprise. The study also evaluates risks across twenty harnesses, highlights vendor dependence, and proposes an internal control plane for future harness ownership decisions.

arXiv AI
Jul 9

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

arXiv:2607. 06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value.

By Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh
arXiv AI
Aug 24

Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work

The paper discusses how large enterprises can adopt the harness paradigm to overcome limitations of traditional coding approaches. It reviews recent findings that harnesses outperform complex architectures at the task level, that harness choice drives benchmark variance more than model choice, and that governance is the main barrier to enterprise adoption. The authors propose a unified harness architecture that keeps code identical across deployments, simplifying review and audit processes.

By George Juraj Salapa
arXiv Computation and Language
Sep 4

Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

The paper introduces HARNESSEVO, a method that decomposes a large language model’s harness into four independently evolvable components—role, task‑strategy, tool/format‑rules, and reflection/control. Experiments on ALFWorld show that overall success rates are similar to flat‑string evolution, but the reflection/control component alone accounts for most of the performance gains. The study also finds that evenly distributing optimization budget across all slots can be detrimental; concentrating resources on the high‑credit control slot recovers lost performance, while on WebShop all slots remain ineffective, suggesting task‑specific differences in harness value.

By Michael Nguyen, Wei Chen Tan, Nurul Aisyah Hassan, Arvind Raman, Li Hua Lim, Ahmad Faiz Razak
arXiv AI
Sep 7

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

The paper introduces $ au^ au$-Bench, a benchmark that turns the construction of AI agents into a measurable task. In this environment a developer agent receives real business records, client requirements, a production API, an existing codebase, and constraints on cost and models, and must deliver a complete customer‑service agent. The benchmark evaluates performance by deploying the agent against simulated users, revealing that current state‑of‑the‑art models achieve only 23.9% success while an expert‑written reference scores 82.2%.

By Quan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres
arXiv AI
Sep 10

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

arXiv:2609.09134v1 Announce Type: new Abstract: Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic...

By Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur
arXiv AI
Aug 28

Five Primitives for Governing Autonomous AI Agents at Runtime

The paper proposes five runtime primitives—discovery, identity, governance, attestation, and supply chain—to manage autonomous AI agents in enterprise settings. It argues that traditional control models fail because agents are transient, model-driven, and self‑discoverable, making runtime governance essential. The authors detail an implementation that mediates agent actions against policy, authorizes them via a per‑tenant vocabulary, and records them in a verifiable ledger, noting the associated operational costs and partial deployment status.

By Jiten Oswal, John Cadeddu
arXiv Computation and Language
Sep 2

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

arXiv:2609.01437v1 Announce Type: cross Abstract: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly...

By Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang
arXiv AI
2d ago

AX is the New AEO

arXiv:2609.34951v2 Announce Type: replace Abstract: In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training...

By Ido Finder, Assaf Elovic, Gad Shalev, Liad Yosef
arXiv Machine Learning
Sep 1

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.

By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu