AEGIS: Runtime-Guided GPU Collocation for Multi-Tenant Deep Learning Training
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 09200v1 Announce Type: cross Abstract: The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems.
The paper introduces COMPASS-ABS, a scheduling framework for shared GPU clusters that reduces resource fragmentation for deep learning training jobs. It defines a new metric, Scheduler-Induced Fragmentation (SIF), which does not rely on historical workload data, and presents the COMPASS algorithm that confines cluster states within an Anchor-Based Space (ABS) to keep fragmentation low. Experiments on both a physical and a simulated cluster show that COMPASS-ABS improves resource utilization and shortens job completion times by mitigating fragmentation.
arXiv:2609.38090v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-...
arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.
arXiv:2504. 11320v4 Announce Type: replace-cross Abstract: Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day.
arXiv:2607. 01646v2 Announce Type: replace Abstract: State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack.