arXiv AI By Zhixu Du, Weijia Han, Hai Helen Li, Yiran Chen

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

Read the original on arXiv AI →

The paper introduces a Mixture-of-Agents (MoA) approach to quantify the per-token compute required by large language models. By having a panel of fifteen models of varying sizes attempt to reproduce each token, the authors define the smallest successful agent’s inference cost as the token’s sufficient compute, providing an upper bound on necessary computation. Experiments on benchmarks show that a 0.5B model can reproduce most tokens, and that the MoA-derived compute map can reduce latency in model routing and drafting tasks while improving or maintaining accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

The paper presents a systematic study of token consumption in AI agents performing coding tasks. It finds that agentic tasks are far more expensive—about 1000 times more tokens than code reasoning or chat—primarily due to input tokens, and that token usage varies wildly, with accuracy peaking at moderate costs. Models differ significantly in efficiency, and current frontier models cannot reliably predict their own token usage, often underestimating it.

By Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei