arXiv AI By Bochao Feng, Jianjiang Li, Haojie Wang, Lin Qiao, Yinghui Li, Yukun Yan, Jidong Zhai

Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows

Read the original on arXiv AI →

The paper introduces a tail‑risk‑aware scheduling strategy for agentic LLM workflows that decouples readiness from immediate release of model turns. By jointly selecting which ready turn to release and controlling the amount of unfinished work kept in the queue, the method uses a mean‑CVaR objective to adapt to evolving tail risk and online turn‑work estimates. Experiments on real software‑engineering task traces show comparable performance to eager release under light load and a significant reduction in the 95th‑percentile workflow flow time, achieving up to a 3.5× speedup under contention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

TOPAS is a Task‑Oriented Prefix‑Aware Scheduler designed for multi‑agent large language model serving. It jointly decides which agent prefixes to retain in a shared key‑value cache and which requests to schedule, balancing the reduction of each task’s longest remaining service path against the benefit of downstream prefix reuse while accounting for movement and preemption costs. Experiments on synthetic DAGs and MetaGPT software‑development workflows show that TOPAS can reduce mean and p99 job completion times by up to 39.8%/49.4% and 22.0%/26.6% respectively compared to the best baselines.

By Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li, Haisheng Tan
arXiv AI
Sep 3

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

The paper investigates why large language model (LLM) agents fail on long, multi‑step production workflows despite high benchmark success. By testing nine models (1.2 B–671 B parameters) across six task families and multiple horizons, the authors find that task success follows a geometric decay governed by a per‑step reliability that never reaches 1, leading to inevitable collapse for long horizons. The degradation is driven mainly by step count rather than context length, and the study quantifies a significant gap between benchmark and production performance, especially for agentic tool‑use tasks.

By Shubhra Mittal
arXiv AI
Aug 26

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a new benchmark designed to evaluate how large language model agents invoke multiple tools while respecting resource constraints and parallel execution. It provides executable multi‑tool workflows with dependency annotations and measured resource profiles, and introduces a two‑part evaluation framework that separates logical planning from physical scheduling. The study shows that strong logical planning alone does not guarantee safe or efficient execution, and that providing resource information can reduce overflows and improve utilization.

By Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Jul 28

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

arXiv:2607. 23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency.

By Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)