Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
arXiv:2606. 18543v1 Announce Type: new Abstract: Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.
arXiv:2608. 08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work.
The paper introduces $ au^ au$-Bench, a benchmark that turns the construction of AI agents into a measurable task. In this environment a developer agent receives real business records, client requirements, a production API, an existing codebase, and constraints on cost and models, and must deliver a complete customer‑service agent. The benchmark evaluates performance by deploying the agent against simulated users, revealing that current state‑of‑the‑art models achieve only 23.9% success while an expert‑written reference scores 82.2%.
E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
arXiv:2509. 26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions.
arXiv:2606. 16613v1 Announce Type: new Abstract: As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important.
arXiv:2609.37658v1 Announce Type: new Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-te...
The paper discusses how enterprises increasingly deploy AI coding agent harnesses, often purchased from vendors like Anthropic or OpenAI, and how these harnesses dictate model choice, prompt handling, and cost. It introduces a fast, customizable routing system that classifies prompts and strategically routes them to minimize expensive model usage, achieving 14–21% cost savings in a simulated 10,000-seat enterprise. The study also evaluates risks across twenty harnesses, highlights vendor dependence, and proposes an internal control plane for future harness ownership decisions.
FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision‑making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities, set lineups, and respond to a board that can fire it, all while a deterministic engine aggregates the outcomes into a final score without human judgment. The benchmark evaluates six behavioral capabilities and compares 15 frontier models in solo and arena tracks, revealing that managerial behavior—not computational scale—drives performance.
arXiv:2606. 15862v1 Announce Type: new Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
arXiv:2603. 16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments.
Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator...