arXiv Machine Learning
Sep 14

MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling

MCRL2 is a reinforcement learning framework that enhances microservice scheduling in cloud data centers by integrating multi-resource cross-attention-based representation learning. It introduces MCRL, a representation learning component that captures structured interactions among nodes, resources, and microservices, and couples this with an actor‑critic architecture and a maximum entropy objective. Experiments on real production cluster traces show that MCRL2 outperforms existing baselines in load balancing, scheduling success rate, and average completion time across diverse workloads.

By Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao
arXiv AI
Sep 4

PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

PPO-STGNN is a DAG task‑scheduling algorithm that combines proximal policy optimization with spatio‑temporal graph neural networks. It extracts features from both the task topology and the heterogeneous cloud‑edge‑end resource graph, then optimizes scheduling to reduce makespan and schedule length ratio while balancing CPU and memory loads. A multi‑teacher behavior‑cloning pretraining step accelerates convergence, and experiments show significant load‑balancing improvements with low completion times in dynamic, heterogeneous environments.

By Yangshuo Qi, Chenwei Wang, Zihan Shen, Songlin Sun
Hugging Face Trending Papers
Sep 3

PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

PPO-STGNN is a DAG task‑scheduling algorithm that combines proximal policy optimization with spatio‑temporal graph neural networks to address the NP‑hard scheduling problem in heterogeneous cloud‑edge‑end environments. It extracts features from both the DAG task topology and the physical resource graph, then optimizes the scheduling policy to minimize makespan and schedule length ratio while improving CPU and memory load balancing. A multi‑teacher behavior‑cloning pretraining step accelerates convergence, and experiments demonstrate significant load‑balancing gains with low completion times in dynamic, heterogeneous settings.