arXiv AI By Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao, Pedro Silvestre, Guo Li, Huanzhou Zhu, Llu\'is Vilanova, Peter Pietzuch

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Read the original on arXiv AI →

Scepsy is a serving system designed to efficiently schedule arbitrary multi‑LLM agentic workflows on GPU clusters. It leverages the observation that each LLM’s share of execution time remains relatively stable across requests, profiling LLMs under various parallelism levels to build an Aggregate LLM Pipeline that predicts throughput and latency. Using this predictor, Scepsy searches for optimal GPU allocations—balancing fractional GPU shares, tensor parallelism, and replica counts—and then heuristically places them on the cluster to reduce fragmentation and honor network topology, achieving up to 2.5× higher throughput and 1.0–3.3× lower latency compared to baseline approaches.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

TOPAS is a Task‑Oriented Prefix‑Aware Scheduler designed for multi‑agent large language model serving. It jointly decides which agent prefixes to retain in a shared key‑value cache and which requests to schedule, balancing the reduction of each task’s longest remaining service path against the benefit of downstream prefix reuse while accounting for movement and preemption costs. Experiments on synthetic DAGs and MetaGPT software‑development workflows show that TOPAS can reduce mean and p99 job completion times by up to 39.8%/49.4% and 22.0%/26.6% respectively compared to the best baselines.

By Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li, Haisheng Tan
arXiv AI
Sep 10

Diamond Agent: Agentic Control of Federated HPC Resources as a Service

arXiv:2609.06181v1 Announce Type: cross Abstract: Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving...

By Haotian Xie, Junlin Chen, Mingkai Zheng, Yifan Zhu, Minu Mathew, Max Burnette, Yadu Babuji, Volodymyr Kindratenko, Shivaram Venkataraman, Kyle Chard, Ian Foster, Zhao Zhang