arXiv AI

CEO-Bench: Can Agents Play the Long Game?

arXiv:2606. 18543v1 Announce Type: new Abstract: Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.

arXiv AI
Sep 7

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

The paper introduces $ au^ au$-Bench, a benchmark that turns the construction of AI agents into a measurable task. In this environment a developer agent receives real business records, client requirements, a production API, an existing codebase, and constraints on cost and models, and must deliver a complete customer‑service agent. The benchmark evaluates performance by deploying the agent against simulated users, revealing that current state‑of‑the‑art models achieve only 23.9% success while an expert‑written reference scores 82.2%.

By Quan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres
arXiv Machine Learning
Sep 1

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.

By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv AI
Sep 25

Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise

The paper discusses how enterprises increasingly deploy AI coding agent harnesses, often purchased from vendors like Anthropic or OpenAI, and how these harnesses dictate model choice, prompt handling, and cost. It introduces a fast, customizable routing system that classifies prompts and strategically routes them to minimize expensive model usage, achieving 14–21% cost savings in a simulated 10,000-seat enterprise. The study also evaluates risks across twenty harnesses, highlights vendor dependence, and proposes an internal control plane for future harness ownership decisions.

By Arian Abbasi, Alan Aqrawi, Ted Kwartler
Hugging Face Trending Papers
Aug 19

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision‑making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities, set lineups, and respond to a board that can fire it, all while a deterministic engine aggregates the outcomes into a final score without human judgment. The benchmark evaluates six behavioral capabilities and compares 15 frontier models in solo and arena tracks, revealing that managerial behavior—not computational scale—drives performance.