arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.
By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv:2602. 09345v3 Announce Type: replace-cross Abstract: AI agents are increasingly deployed in multi-tenant cloud environments, where they execute diverse tool calls within sandboxed containers, each call with distinct resource demands and rapid fluctuations.
By Yusheng Zheng, Jiakun Fan, Quanzhi Fu, Yiwei Yang, Wei Zhang, Andi Quinn
The paper investigates how large‑language‑model (LLM) based AI agents mix latency, local resource usage, and container bottlenecks when processing user requests that involve remote LLM calls and local tool execution. By measuring three representative tasks—retrieval‑augmented question answering, web search, and software coding—the authors show that agents exhibit diverse resource dynamics, with concurrent requests revealing task‑specific bottlenecks in CPU, disk I/O, and memory. Leveraging these insights, they propose CPU‑aware tool admission and task‑aware CPU allocation, achieving up to a 5.4× speed‑up for CPU‑sensitive tasks and a 32% reduction in average latency across multiple tasks.
By Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang
Scepsy is a serving system designed to efficiently schedule arbitrary multi‑LLM agentic workflows on GPU clusters. It leverages the observation that each LLM’s share of execution time remains relatively stable across requests, profiling LLMs under various parallelism levels to build an Aggregate LLM Pipeline that predicts throughput and latency. Using this predictor, Scepsy searches for optimal GPU allocations—balancing fractional GPU shares, tensor parallelism, and replica counts—and then heuristically places them on the cluster to reduce fragmentation and honor network topology, achieving up to 2.5× higher throughput and 1.0–3.3× lower latency compared to baseline approaches.
By Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao, Pedro Silvestre, Guo Li, Huanzhou Zhu, Llu\'is Vilanova, Peter Pietzuch
arXiv:2604. 26963v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as the execution core of autonomous agents rather than as standalone text generators.
By Yifei Wang, Hancheng Ye, Yechen Xu, Cong Guo, Chiyue Wei, Qinsi Wang, Dongting Li, Tingjun Chen, Hai "Helen" Li, Danyang Zhuo, Yiran Chen
arXiv:2609.06181v1 Announce Type: cross
Abstract: Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving...
By Haotian Xie, Junlin Chen, Mingkai Zheng, Yifan Zhu, Minu Mathew, Max Burnette, Yadu Babuji, Volodymyr Kindratenko, Shivaram Venkataraman, Kyle Chard, Ian Foster, Zhao Zhang
arXiv:2608. 00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different from chatbots.
By Banruo Liu, Haoran Qiu, \'I\~nigo Goiri, Rodrigo Fonseca, Ricardo Bianchini, Esha Choukse
arXiv:2607. 23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency.
By Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)
arXiv:2606. 11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features.
By Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou
The paper introduces agentic-eCAL, an extension of the Energy Cost of AI Lifecycle metric to evaluate multi‑agent AI workflows across the edge‑cloud continuum. By combining a two‑rate energy model with OSI‑layer transport analysis, the authors quantify that inter‑agent text transfer accounts for only 0.25% of total workflow energy, highlighting that the main energy cost lies in additional inference and context processing triggered by communication. The study uses extensive GPU benchmarks on NVIDIA A100/H100 with 16 open‑weight models and 8 orchestration topologies to validate the metric and explore placement implications.
By Carolina Fortuna, Vid Han\v{z}el, Tim Strnad, Bla\v{z} Bertalani\v{c}
Multi-turn LLM agents interleave model calls with external tool invocations, shifting serving from stateless request processing to stateful program execution. Serving these workloads requires scheduling, KV-cache management, and routing policies that use program-level context, including turn dependencies, tool-induced gaps, and reusable KV state.
arXiv:2606. 09613v1 Announce Type: cross Abstract: Multi-turn LLM agents interleave model calls with external tool invocations, shifting serving from stateless request processing to stateful program execution.
By Rakibul Hasan Rajib, Mengxin Zheng, Qian Lou