arXiv AI

User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

arXiv:2608. 11840v1 Announce Type: cross Abstract: Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand.

arXiv Machine Learning
Jul 22

Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks

arXiv:2607. 18288v1 Announce Type: new Abstract: Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization.

By Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas
arXiv Machine Learning
Sep 14

MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling

MCRL2 is a reinforcement learning framework that enhances microservice scheduling in cloud data centers by integrating multi-resource cross-attention-based representation learning. It introduces MCRL, a representation learning component that captures structured interactions among nodes, resources, and microservices, and couples this with an actor‑critic architecture and a maximum entropy objective. Experiments on real production cluster traces show that MCRL2 outperforms existing baselines in load balancing, scheduling success rate, and average completion time across diverse workloads.

By Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao
arXiv AI
3d ago

Inference Auctions

The paper proposes an inference auction for large language model (LLM) APIs, enabling users to bid for faster service when compute demand exceeds capacity. The auction allocates priority efficiently without increasing latency, and includes fast algorithms for truthful bidding and an autobidding agent that adjusts bids within a user’s budget to maximize utility. Experiments show the auction improves system welfare while preserving the cache utilization and latency benefits of the SGLang inference framework.

By Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan
arXiv Machine Learning
Jul 8

Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach

arXiv:2605. 02965v2 Announce Type: replace Abstract: Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content, giving rise to rapidly growing computational workloads in cloud data centers.

By Yang Fu, Peng Qin, Liming Chen, Zihao Zhang, Hao Yu, Yifei Wang
arXiv Machine Learning
Aug 31

DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge

DART-FL is a multitask federated learning framework designed for edge devices that must balance online inference and model training under limited resources. It dynamically allocates resources between inference and training based on current inference backlog and service capacity, then distributes remaining training capacity among tasks using a queue‑aware scheduler that adjusts loss weights. Experiments on image classification datasets with synthetic and real workloads show that DART‑FL adapts to bursty inference demand, improving accuracy for high‑demand tasks while preserving overall multitask performance.

By Yiming Xie, Pinrui Yu, Geng Yuan, Xue Lin, Ningfang Mi