The paper introduces Global Clustered Parallel Split Learning (GCPSL), a method that groups clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that distributing the population across more workloads increases direct data participation, though smaller clusters may slightly reduce accuracy. In a practical four‑GPU setup, label‑aware GCPSL achieves 85 % CIFAR‑10 validation accuracy in roughly 6 minutes, compared to over 19 minutes for serialized workloads, with size‑balanced cluster assignments improving participation by 3.25 percentage points.
By Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink
CluSTER is a cluster‑aware balanced sampling framework designed to improve the efficiency of instruction‑tuning for large language models. It reduces redundant computation by clustering data in gradient space and allocating samples across GPUs in a data‑parallel setting, while preserving the original distribution through weighted updates. Experiments show that CluSTER can cut training time by up to 69.6% with negligible loss in accuracy compared to existing sampling methods.
By Hyunjin Kim, Youngeun Nam, Jaemin Han, Wonhyeok Choi, Jae-Gil Lee
Asynchronous federated learning improves scalability by updating the global model from a server-side buffer of client updates as they arrive, rather than waiting for all selected clients to finish. Wh...
arXiv:2609.07192v1 Announce Type: cross
Abstract: Asynchronous federated learning improves scalability by updating the global model from a server-side buffer of client updates as they arrive, rather...
By Prashant Bajpai, Divya Saxena, Philippe Lalanda, German Vega
arXiv:2609.39350v1 Announce Type: cross
Abstract: As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training para...
By Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, Quanlu Zhang, Yu Wang
The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.
By Gianluca Mittone, Marco Aldinucci
The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By forming loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to traditional sharded DP approaches.
arXiv:2606. 01007v1 Announce Type: cross Abstract: Sparsely activated Mixture-of-Experts (MoE) models scale capacity via conditional computation, but distributed inference suffers from cross-GPU expert communication and routing-induced load imbalance.
By Zhiyao Xu, Aoxue Liu, Zhanjie Ding, Dan Zhao, Yong Jiang, Qing Li
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving supe...
arXiv:2606. 00946v1 Announce Type: cross Abstract: Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for GPU utilization.
By Gangmuk Lim, Wanyu Zhao, Brighten Godfrey, Jiaxin Shan, Le Xu, Liguang Xie
arXiv:2607. 01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models.
By Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty
The paper introduces COMPASS-ABS, a scheduling framework for shared GPU clusters that reduces resource fragmentation for deep learning training jobs. It defines a new metric, Scheduler-Induced Fragmentation (SIF), which does not rely on historical workload data, and presents the COMPASS algorithm that confines cluster states within an Anchor-Based Space (ABS) to keep fragmentation low. Experiments on both a physical and a simulated cluster show that COMPASS-ABS improves resource utilization and shortens job completion times by mitigating fragmentation.
By Yukai Zhou, Hongfan Wu