arXiv Machine Learning

COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

The paper introduces COMPASS-ABS, a scheduling framework for shared GPU clusters that reduces resource fragmentation for deep learning training jobs. It defines a new metric, Scheduler-Induced Fragmentation (SIF), which does not rely on historical workload data, and presents the COMPASS algorithm that confines cluster states within an Anchor-Based Space (ABS) to keep fragmentation low. Experiments on both a physical and a simulated cluster show that COMPASS-ABS improves resource utilization and shortens job completion times by mitigating fragmentation.

arXiv Machine Learning
Jun 30

Harvesting AI Computation at the Edge via Generic Approximation

arXiv:2606. 29518v1 Announce Type: cross Abstract: With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge.

By Yihan Wang, Huiru Yan, Luxin Zhang, Long Cheng, Weiwei Chen, Ying Wang, Lei Zhang, Cheng Liu, Huawei Li
Hugging Face Trending Papers
Sep 24

Concurrent Split Learning Through Stable Client Clustering

The paper introduces Global Clustered Parallel Split Learning (GCPSL), which partitions clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that increasing the number of concurrent workloads boosts direct data participation, though smaller clusters may slightly reduce accuracy. On a four‑GPU setup, label‑aware GCPSL achieves 85% CIFAR‑10 validation accuracy in about 6.13 minutes, compared to 19.09 minutes for serialized workloads, and size‑balanced cluster assignments improve participation by 3.25 percentage points.