arXiv AI

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

The paper introduces the concept of "bottling"—the ability of large language model (LLM) agents to transform general capabilities into task‑specific, cost‑effective solutions for large, repetitive workloads. It presents BOTTLED, a benchmark where agents receive an unlabelled workload and must complete it within fixed time, compute, and API budgets, choosing strategies such as training small models or writing reusable programs. Experiments across ten models and three tasks show that strong zero‑shot performance does not guarantee effective bottling, yet bottling can still achieve substantial cost savings and competitive performance compared to specialized cheap inference models.

arXiv AI
Jul 28

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

arXiv:2607. 23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes.

By Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
arXiv AI
Sep 24

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

The paper introduces the concept of an agent’s "taste"—its ability to make effective long‑horizon decisions—and presents Taste‑Bench, a new benchmark that automatically generates decision‑fork questions from agent trajectories. Taste‑Bench evaluates models on choosing the best path without seeing future outcomes, revealing that top models answer only about 60% of questions correctly and that later‑appearing evidence makes forks harder. The authors also demonstrate that training a student model to mimic a teacher’s judgment improves decision quality and overall success on held‑out software engineering tasks.

By Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang, Yan Lu, Qi He, Xiaohua Jia
arXiv Computation and Language
Sep 3

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

EarlyEval introduces a lightweight framework that predicts an LLM agent’s final outcome early in its execution, allowing the run to halt when a LightGBM classifier reaches a calibrated confidence threshold. By training success and failure classifiers on behavioral, textual, and reference-solution features, EarlyEval can cut 13%-26% of agent steps and up to 44.1% of input tokens while maintaining 89%-97% prediction accuracy. Across three benchmarks—SWE-bench Verified, TerminalBench, and Toolathlon—this approach reduces evaluation costs with minimal impact on per-agent resolve rates.

By Yuling Shi, Zhensu Sun, Junsen Dong, Chengcheng Wan, David Lo, Xiaodong Gu
arXiv AI
Sep 3

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

The paper introduces SCX Router, a lightweight GLiClass-based model selector that assigns suitability scores to inference-time language models without autoregressive generation. It uses a 0.6B-parameter Qwen3 decoder with a shallow bidirectional scorer, preserving a text-only key–value cache across sessions and predicting task attributes such as type, difficulty, and expected output length. The authors build a comprehensive task ontology with 23 families, 115 types, and 1,173 synthetic examples, generating 150,000 verifier-scored tasks to train the router, which outperforms baseline models on LiveBench subsets with a top‑1 score of 0.707 versus 0.696 for the strongest fixed model.

By Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov
arXiv Computer Vision
Aug 28

OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks

OS-Marathon is a new benchmark that tests computer‑use agents on vast‑horizon, repetitive tasks, covering 100 tasks across five scenarios and ten domains. The study shows that current state‑of‑the‑art agents perform poorly on these tasks, and that simply decomposing workflows into subtasks does not solve the problem. Introducing a cost‑friendly personalization method called GraphDemo, which adapts agents from a single human demonstration, improves performance, highlighting the value of human guidance for these challenging tasks.

By Jing Wu, Wenjie Ai, Daphne Barretto, Yiye Chen, Qingyu Chen, Yuhang He, Pranit Chawla, Nicholas Gyd\'e, Yanan Jian, Vibhav Vineet
arXiv Machine Learning
Sep 18

Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement

The paper introduces Poodle, a prototype for just‑in‑time model replacement (JITR) that automatically swaps a large language model with a cheaper, task‑specific model when a recurring task is detected. Poodle reduces inference time by up to 7.5× and saves over $2,200 per 1 M requests compared to a flagship hosted LLM, while maintaining competitive accuracy. The authors argue that model search and transfer learning are essential for efficiently identifying and fine‑tuning these custom models.

By Nils Strassenburg, Boris Glavic, Tilmann Rabl
arXiv AI
1d ago

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS is a method that builds reusable memory for large‑language‑model agents by having an explorer agent generate self‑created tasks and a solver agent attempt them. When the solver fails, a heuristic is extracted and only accepted after repeated successful use, then added to a memory bank for future test‑time use. Experiments on AppWorld, τ²‑bench, and AutomationBench show that DAEDALUS raises mean success rates by up to 15.9 points and pass⁵ by up to 2.2× compared to a no‑memory baseline, while also providing a cost‑effective alternative to training‑task or oracle‑verifier approaches.

By Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud