arXiv:2609.14500v1 Announce Type: new
Abstract: AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a hi...
By Seyed Morteza Emadi
arXiv:2607. 13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires.
By Junjie Yin, Xinyu Feng
The paper introduces the concept of substrate blindness, where AI agents lack execution context in their planning. By providing a 128 MB RAM and 10 s wall‑time contract to large language models, the authors show that agents generate code that uses less memory, runs faster, and incorporates structural changes such as bounded blocking and in‑place buffers. Across three leading models, contract disclosure improved resource usage and correctness, demonstrating that minimal execution contracts can guide agents to produce more efficient programs.
By Manu Agrawal
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway.
arXiv:2609.23058v1 Announce Type: new
Abstract: Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for...
By Xin Heng
arXiv:2607. 28666v1 Announce Type: cross Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained.
By Prerit Ahuja
arXiv:2606. 04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks.
By Jingbo Wen, Liang He, Ziqi He
The paper investigates how agent harnesses—specifically planning guidance, execution organization, and completion verification—affect performance in retail and airline pilot tasks. By comparing fixed, task‑specific plans to shuffled policy text of equal length, the study finds that fixed plans improve success rates by about 7 percentage points, especially on complex tasks. A read‑only verifier rejects a majority of invalid episodes while incurring minimal cost, and its impact varies with the penalty for erroneous acceptance, often matching the full planning‑plus‑verification benefit at a lower cost.
By Yukun Zhang, Kemu Xu, Yishen Chen
arXiv:2607. 06906v1 Announce Type: new Abstract: Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value.
By Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh
arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
By Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
arXiv:2608. 07474v1 Announce Type: new Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max.
By Hiroki Naito
arXiv:2606. 17930v1 Announce Type: new Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving.
By Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec