MCRL2 is a reinforcement learning framework that enhances microservice scheduling in cloud data centers by integrating multi-resource cross-attention-based representation learning. It introduces MCRL, a representation learning component that captures structured interactions among nodes, resources, and microservices, and couples this with an actor‑critic architecture and a maximum entropy objective. Experiments on real production cluster traces show that MCRL2 outperforms existing baselines in load balancing, scheduling success rate, and average completion time across diverse workloads.
By Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao
arXiv:2606. 11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features.
By Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou
arXiv:2607. 05272v1 Announce Type: cross Abstract: Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic.
By Ruslan Sharifullin
arXiv:2606. 07565v1 Announce Type: new Abstract: Intelligent scaling of microservices in cloud platforms is crucial for mitigating escalating compute costs while avoiding service disruptions.
By Ahmed Abdulaal, Maruf Aytekin, Thilaga kumaran Srinivasan, Tomer Lancewicki
The paper introduces a stability‑aware autoscaling framework for edge serverless workloads that combines an Attention‑Enhanced Double‑Stacked LSTM with Proximal Policy Optimization to address temporal blindness in deep reinforcement learning. By weighting recent historical states non‑uniformly, the method suppresses high‑frequency jitter while preserving demand trends, outperforming single‑layer LSTM, static HPA, and KEDA baselines in latency reduction and stability. Experiments on two Kubernetes clusters with real Azure Functions traces show a ~67% reduction in P90 latency and improved adherence to a 50 ms hard SLO.
By Faraz Shaikh, Gianluca Reali, Mauro Femminella
Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic. We investigate whether reinforcement learning (RL) can learn adaptive batching and routing policies that outperform these heuristics, training REINFORCE and PPO agents on a discrete-event simulator validated against queuing theory and production traces (Azure Functions, BurstGPT).
The paper introduces a taxonomy for resource management systems that span IoT, edge, and cloud layers, extending Wang et al.’s Continuum Orchestration Systems with two new dimensions: the AI Augmentation Paradigm and the Feedback channel. It evaluates six recent architectures and finds that none combine full LLM orchestration with full agent‑layer feedback in a Cloud Continuum setting, highlighting a gap in cross‑tier feedback abstraction. The study emphasizes the need for a unified feedback mechanism to bridge disparate per‑tier signals to the LLM orchestrator.
By Antonino Vaccarella, Lanpei Li, Vincenzo Lomonaco, Massimo Coppola
arXiv:2609.14952v1 Announce Type: new
Abstract: Dynamic cloud workflow scheduling must balance deadline satisfaction, container utilization, and energy consumption while dealing with stochastic task-...
By Zongjin Li, Shaohan Feng, Chunxi Yang, Wenbo Wang
arXiv:2609.14968v1 Announce Type: new
Abstract: Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay betwe...
By Tiangang Li, Shi Ying, Xiangbo Tian
Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.
By Hamed Hamzeh
arXiv:2608. 11840v1 Announce Type: cross Abstract: Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand.
By Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magn\'usson, Praveen Kumar Donta
arXiv:2607. 18288v1 Announce Type: new Abstract: Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization.
By Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas