arXiv AI By Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Read the original on arXiv AI →

The paper presents a systematic study of token consumption in AI agents performing coding tasks. It finds that agentic tasks are far more expensive—about 1000 times more tokens than code reasoning or chat—primarily due to input tokens, and that token usage varies wildly, with accuracy peaking at moderate costs. Models differ significantly in efficiency, and current frontier models cannot reliably predict their own token usage, often underestimating it.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

The paper introduces a Mixture-of-Agents (MoA) approach to quantify the per-token compute required by large language models. By having a panel of fifteen models of varying sizes attempt to reproduce each token, the authors define the smallest successful agent’s inference cost as the token’s sufficient compute, providing an upper bound on necessary computation. Experiments on benchmarks show that a 0.5B model can reproduce most tokens, and that the MoA-derived compute map can reduce latency in model routing and drafting tasks while improving or maintaining accuracy.

By Zhixu Du, Weijia Han, Hai Helen Li, Yiran Chen
arXiv AI
Aug 28

Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

The paper argues that large language models need adaptive reasoning rather than fixed reasoning budgets. It shows that over‑reasoning leads to high computational cost without accuracy gains, while under‑reasoning results in incorrect or incomplete solutions. The authors evaluate these failure modes on MATH‑500 and the GAIA benchmark, highlighting the need for dynamic reasoning allocation in agentic AI systems.

By Md Jueal Mia, M. Hadi Amini
arXiv AI
Sep 18

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

The paper investigates how large‑language‑model (LLM) based AI agents mix latency, local resource usage, and container bottlenecks when processing user requests that involve remote LLM calls and local tool execution. By measuring three representative tasks—retrieval‑augmented question answering, web search, and software coding—the authors show that agents exhibit diverse resource dynamics, with concurrent requests revealing task‑specific bottlenecks in CPU, disk I/O, and memory. Leveraging these insights, they propose CPU‑aware tool admission and task‑aware CPU allocation, achieving up to a 5.4× speed‑up for CPU‑sensitive tasks and a 32% reduction in average latency across multiple tasks.

By Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang