Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway.
The paper investigates how the distribution of capacity and task information across stages of AI production affects the economic value of inference. Using controlled workflow experiments on software‑engineering tasks, it finds that direct execution achieves a 59.6% success rate at token ceilings of 12,000 and 24,000, while information‑constrained planning improves from 36.2% to 51.2%. The study also shows that giving planners access to task issues boosts success, and that scaling token limits changes the balance between planning and execution workloads.
By Yukun Zhang, Kemu Xu, Yishen Chen
arXiv:2607. 06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value.
By Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh
Budgets for AI tokens can’t be infinite, no matter how much hyperscalers wish they were The post Drilling Into AI’s Financial Sustainability appeared first on Towards Data Science .
By Stephanie Kirmer
arXiv:2609.35760v2 Announce Type: replace-cross
Abstract: When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The age...
By Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
arXiv:2609.33965v2 Announce Type: replace-cross
Abstract: We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between i...
By Joshua Horswill, Ross Hunter, Matt Clifford, James Hall
Google commits $40M in AI tokens and credits for the Genesis Mission
arXiv:2606. 14769v1 Announce Type: cross Abstract: Agentic AI systems are increasingly being deployed as productive resources in organizational workflows, yet existing evaluation methods primarily measure isolated technical performance rather than economic contribution.
By Quanyan Zhu
arXiv:2606. 07632v1 Announce Type: new Abstract: Proper accounting of the energy requirements and environmental impact of artificial intelligence (AI) systems is necessary for researchers, developers, policy makers, and users to assess the barriers to building systems at scale.
By Jared Fernandez, Clara Na, Yonatan Bisk, Constantine Samaras, Emma Strubell
arXiv:2606. 08998v3 Announce Type: replace Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer.
By Muhammad Zia Hydari, Raja Iqbal
arXiv:2606. 08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer.
By Muhammad Zia Hydari, Raja Iqbal
arXiv:2606. 26118v1 Announce Type: cross Abstract: We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations.
By Seamus Somerstep, Aritra Guha, Divesh Srivastava, Yuekai Sun