The paper introduces Global Clustered Parallel Split Learning (GCPSL), which partitions clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that increasing the number of concurrent workloads boosts direct data participation, though smaller clusters may slightly reduce accuracy. On a four‑GPU setup, label‑aware GCPSL achieves 85% CIFAR‑10 validation accuracy in about 6.13 minutes, compared to 19.09 minutes for serialized workloads, and size‑balanced cluster assignments improve participation by 3.25 percentage points.
arXiv:2609.07192v1 Announce Type: cross
Abstract: Asynchronous federated learning improves scalability by updating the global model from a server-side buffer of client updates as they arrive, rather...
By Prashant Bajpai, Divya Saxena, Philippe Lalanda, German Vega
Asynchronous federated learning improves scalability by updating the global model from a server-side buffer of client updates as they arrive, rather than waiting for all selected clients to finish. Wh...
CluSTER is a cluster‑aware balanced sampling framework designed to improve the efficiency of instruction‑tuning for large language models. It reduces redundant computation by clustering data in gradient space and allocating samples across GPUs in a data‑parallel setting, while preserving the original distribution through weighted updates. Experiments show that CluSTER can cut training time by up to 69.6% with negligible loss in accuracy compared to existing sampling methods.
By Hyunjin Kim, Youngeun Nam, Jaemin Han, Wonhyeok Choi, Jae-Gil Lee
The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.
By Gianluca Mittone, Marco Aldinucci
arXiv:2608. 15639v1 Announce Type: cross Abstract: \textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients.
By Wenhao Yuan, Chenchen Lin, Wenhao Hu, Jian Chen, Jinfeng Xu, Shujie Li, Edith Cheuk Han Ngai
The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By forming loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to traditional sharded DP approaches.
arXiv:2606. 01007v1 Announce Type: cross Abstract: Sparsely activated Mixture-of-Experts (MoE) models scale capacity via conditional computation, but distributed inference suffers from cross-GPU expert communication and routing-induced load imbalance.
By Zhiyao Xu, Aoxue Liu, Zhanjie Ding, Dan Zhao, Yong Jiang, Qing Li
arXiv:2609.39350v1 Announce Type: cross
Abstract: As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training para...
By Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, Quanlu Zhang, Yu Wang
arXiv:2606. 00946v1 Announce Type: cross Abstract: Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for GPU utilization.
By Gangmuk Lim, Wanyu Zhao, Brighten Godfrey, Jiaxin Shan, Le Xu, Liguang Xie
\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split stra...
StoCFL is a clustered federated learning framework designed to address Non-IID data and dynamic client participation. It introduces a flexible clustering mechanism that allows arbitrary client participation and accommodates newly joined clients, improving data efficiency and model performance. Experiments on four Non-IID settings and a real-world dataset demonstrate that StoCFL achieves promising cluster results even when the number of clusters is unknown, outperforming baseline approaches across various scenarios.
By Dun Zeng, Xiangjing Hu, Shiyu Liu, Yue Yu, Qifan Wang, Zenglin Xu