arXiv:2609.14968v1 Announce Type: new
Abstract: Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay betwe...
By Tiangang Li, Shi Ying, Xiangbo Tian
arXiv:2606. 01162v1 Announce Type: new Abstract: Workflow scheduling in cloud computing demands the intelligent allocation of dynamically arriving, graph-structured workflows with varying deadlines onto ever-changing virtual machine resources.
By Ya Shen, Gang Chen, Hui Ma, Mengjie Zhang
MCRL2 is a reinforcement learning framework that enhances microservice scheduling in cloud data centers by integrating multi-resource cross-attention-based representation learning. It introduces MCRL, a representation learning component that captures structured interactions among nodes, resources, and microservices, and couples this with an actor‑critic architecture and a maximum entropy objective. Experiments on real production cluster traces show that MCRL2 outperforms existing baselines in load balancing, scheduling success rate, and average completion time across diverse workloads.
By Tiangang Li, Shi Ying, Xiangbo Tian, Chuan Shi, Ding Xiao
PPO-STGNN is a DAG task‑scheduling algorithm that combines proximal policy optimization with spatio‑temporal graph neural networks. It extracts features from both the task topology and the heterogeneous cloud‑edge‑end resource graph, then optimizes scheduling to reduce makespan and schedule length ratio while balancing CPU and memory loads. A multi‑teacher behavior‑cloning pretraining step accelerates convergence, and experiments show significant load‑balancing improvements with low completion times in dynamic, heterogeneous environments.
By Yangshuo Qi, Chenwei Wang, Zihan Shen, Songlin Sun
arXiv:2606. 11440v1 Announce Type: new Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features.
By Ahasan Kabir, Jiaqi Xue, Mengxin Zheng, Qian Lou
PPO-STGNN is a DAG task‑scheduling algorithm that combines proximal policy optimization with spatio‑temporal graph neural networks to address the NP‑hard scheduling problem in heterogeneous cloud‑edge‑end environments. It extracts features from both the DAG task topology and the physical resource graph, then optimizes the scheduling policy to minimize makespan and schedule length ratio while improving CPU and memory load balancing. A multi‑teacher behavior‑cloning pretraining step accelerates convergence, and experiments demonstrate significant load‑balancing gains with low completion times in dynamic, heterogeneous settings.
arXiv:2603. 23249v2 Announce Type: replace-cross Abstract: Efficient scheduling of directed acyclic graphs (DAGs) is a core problem in large-scale data-intensive computing systems, where query plans, data-processing workloads, and computation graphs consist of dependent tasks competing for limited heterogeneous resource pools.
By Ruisong Zhou, Haijun Zou, Li Zhou, Chumin Sun, Zaiwen Wen
arXiv:2607. 05272v1 Announce Type: cross Abstract: Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic.
By Ruslan Sharifullin
arXiv:2607. 18288v1 Announce Type: new Abstract: Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization.
By Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas
Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.
By Hamed Hamzeh
arXiv:2605. 13221v2 Announce Type: replace Abstract: In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC).
By Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou, Malcolm Yoke Hean Low
arXiv:2609.37085v1 Announce Type: cross
Abstract: Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited a...
By Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar