arXiv AI

SquidAgent: Parallelize Wisely, Coordinate Efficiently

arXiv AI
Sep 7

Substrate-Aware AI Agents: Execution Context as a First-Class Input

The paper introduces the concept of substrate blindness, where AI agents lack execution context in their planning. By providing a 128 MB RAM and 10 s wall‑time contract to large language models, the authors show that agents generate code that uses less memory, runs faster, and incorporates structural changes such as bounded blocking and in‑place buffers. Across three leading models, contract disclosure improved resource usage and correctness, demonstrating that minimal execution contracts can guide agents to produce more efficient programs.

By Manu Agrawal
arXiv AI
Aug 26

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a new benchmark designed to evaluate how large language model agents invoke multiple tools while respecting resource constraints and parallel execution. It provides executable multi‑tool workflows with dependency annotations and measured resource profiles, and introduces a two‑part evaluation framework that separates logical planning from physical scheduling. The study shows that strong logical planning alone does not guarantee safe or efficient execution, and that providing resource information can reduce overflows and improve utilization.

By Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Sep 18

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

The paper investigates how large‑language‑model (LLM) based AI agents mix latency, local resource usage, and container bottlenecks when processing user requests that involve remote LLM calls and local tool execution. By measuring three representative tasks—retrieval‑augmented question answering, web search, and software coding—the authors show that agents exhibit diverse resource dynamics, with concurrent requests revealing task‑specific bottlenecks in CPU, disk I/O, and memory. Leveraging these insights, they propose CPU‑aware tool admission and task‑aware CPU allocation, achieving up to a 5.4× speed‑up for CPU‑sensitive tasks and a 32% reduction in average latency across multiple tasks.

By Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang
arXiv AI
1d ago

Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner

The paper introduces NP‑Bench, a benchmark and a proactive scheduling planner that coordinates parallel large‑language‑model coding agents. By partitioning work scopes and ordering merges ahead of time, the planner improves clean‑integration rates from 1/9 to 9/9 and eliminates merge conflicts, outperforming both no‑coordination and reactive‑detection baselines. It also demonstrates that cross‑session memory can eliminate repeated mistakes and that routing facts to agents does not improve long‑context accuracy at scale.

By Sumanyu Muku
arXiv AI
Sep 10

AgentServeSim: Serving-System Simulation and Policy Search for LLM Agent Programs

AgentServeSim is a simulation framework designed to model the execution of large language model (LLM) agent programs, capturing cross‑turn key‑value (KV) state retention, successor turn release, and scheduling decisions. Unlike existing simulators that operate on request streams, AgentServeSim treats the entire agent program as a single unit of execution, using a Program Control Block, Program Orchestrator, Retention Plane, and Dispatch Plane to emulate realistic serving dynamics. Validation against real vLLM deployments on two GPU platforms shows mean job completion time errors below 5.5%, and the simulator enables automated policy search that improves mean JCT by up to 2.8% over hand‑written policies. whyItMatters":"The simulator provides a realistic, CPU‑based tool for evaluating and optimizing LLM agent serving policies, achieving high fidelity to real deployments and enabling measurable performance gains."

By Rakibul Hasan Rajib, Mengxin Zheng, Qian Lou