COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads
Read the original on arXiv Machine Learning →The paper introduces COMPASS-ABS, a scheduling framework for shared GPU clusters that reduces resource fragmentation for deep learning training jobs. It defines a new metric, Scheduler-Induced Fragmentation (SIF), which does not rely on historical workload data, and presents the COMPASS algorithm that confines cluster states within an Anchor-Based Space (ABS) to keep fragmentation low. Experiments on both a physical and a simulated cluster show that COMPASS-ABS improves resource utilization and shortens job completion times by mitigating fragmentation.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.