arXiv:2608. 06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges.
By Guofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou
arXiv:2402. 09589v2 Announce Type: replace-cross Abstract: We present MLCC, a novel technique to augment today's congestion control algorithms to accelerate DNN training jobs in shared GPU clusters in a fully distributed manner.
By Anton A. Zabreyko, Sanjoli Narang, Sudarsanan Rajasekaran, Manya Ghobadi
arXiv:2608. 15762v1 Announce Type: cross Abstract: Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited.
By Weinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang
arXiv:2606. 29518v1 Announce Type: cross Abstract: With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge.
By Yihan Wang, Huiru Yan, Luxin Zhang, Long Cheng, Weiwei Chen, Ying Wang, Lei Zhang, Cheng Liu, Huawei Li
arXiv:2512. 10236v2 Announce Type: replace-cross Abstract: Modern ML workloads demand distributing training and inference across multiple GPUs.
By Shagnik Pal, Shaizeen Aga, Suchita Pati, Mahzabeen Islam, Lizy K. John
arXiv:2609.14968v1 Announce Type: new
Abstract: Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay betwe...
By Tiangang Li, Shi Ying, Xiangbo Tian
arXiv:2609.39481v1 Announce Type: cross
Abstract: Efficient resource provisioning for large-scale workflows on cloud infrastructures is a critical performance engineering challenge. These workflows a...
By Max Otto, Haci Ismail Aslan, Joel Witzke, Jonathan Bader, Odej Kao
arXiv:2605. 23247v2 Announce Type: replace Abstract: In this paper, we introduce the first machine learning framework for predicting optimal processing times in Single-Level Tree Network (SLTN) architectures for the Divisible Load Theory (DLT) paradigm.
By Bharadwaj Veeravalli
arXiv:2606. 01007v1 Announce Type: cross Abstract: Sparsely activated Mixture-of-Experts (MoE) models scale capacity via conditional computation, but distributed inference suffers from cross-GPU expert communication and routing-induced load imbalance.
By Zhiyao Xu, Aoxue Liu, Zhanjie Ding, Dan Zhao, Yong Jiang, Qing Li
arXiv:2606. 01680v1 Announce Type: cross Abstract: Network failures are among the most frequent hardware faults in large-scale GPU clusters and a leading cause of training-job interruptions.
By Peiqing Chen, Jiedong Jiang, Nengneng Yu, Yuefeng Wang, Sixian Xiong, Wei Wang, Zaoxing Liu
The paper introduces Global Clustered Parallel Split Learning (GCPSL), which partitions clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that increasing the number of concurrent workloads boosts direct data participation, though smaller clusters may slightly reduce accuracy. On a four‑GPU setup, label‑aware GCPSL achieves 85% CIFAR‑10 validation accuracy in about 6.13 minutes, compared to 19.09 minutes for serialized workloads, and size‑balanced cluster assignments improve participation by 3.25 percentage points.
arXiv:2604. 23841v2 Announce Type: replace-cross Abstract: Efficiently solving the Job Shop Scheduling Problem in real-world industrial applications requires policies that are both computationally lean and topologically robust.
By Jonathan Hoss, Moritz Link, Noah Klarmann