Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway.
The paper investigates how the distribution of capacity and task information across stages of AI production affects the economic value of inference. Using controlled workflow experiments on software‑engineering tasks, it finds that direct execution achieves a 59.6% success rate at token ceilings of 12,000 and 24,000, while information‑constrained planning improves from 36.2% to 51.2%. The study also shows that giving planners access to task issues boosts success, and that scaling token limits changes the balance between planning and execution workloads.
By Yukun Zhang, Kemu Xu, Yishen Chen
arXiv:2607. 06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value.
By Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh
Budgets for AI tokens can’t be infinite, no matter how much hyperscalers wish they were The post Drilling Into AI’s Financial Sustainability appeared first on Towards Data Science .
By Stephanie Kirmer
arXiv:2609.35760v2 Announce Type: replace-cross
Abstract: When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The age...
By Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
arXiv:2609.33965v2 Announce Type: replace-cross
Abstract: We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between i...
By Joshua Horswill, Ross Hunter, Matt Clifford, James Hall