The paper introduces MUC-FL, a block‑wise marginal utility contribution framework that selectively transmits only the most impactful data blocks in federated learning to reduce communication overhead. Applied to a multimodal dataset derived from multiple MIMIC clinical datasets, the method identifies 24 out of 1,135 candidate blocks (1.76%) as carrying meaningful improvement signals, potentially cutting communication by 45‑50% while preserving or enhancing model quality. The deduplication‑based block selection achieves a macro F1 score of 0.8566 versus 0.8155 for standard federated optimization, showing improved performance especially for underrepresented classes.
By Akshay Mhatre, Vikram Karthick, Deepti Gupta, Jia Zou
arXiv:2609.39646v1 Announce Type: new
Abstract: Federated learning faces severe communication bottlenecks when clients upload high-dimensional model updates. Existing methods often compress these upd...
By Pengfei Li, Mohammad Khalil
FedSAP is a federated learning framework that addresses heterogeneous edge devices by using structured pruning as a budget-constrained tri-state channel allocation. It partitions model channels into a Global pool, pseudo-domain-specific Private pools, and a Dropped state, allowing broadly useful features to be shared while isolating domain-sensitive updates. Experiments on Digits and Office-Caltech datasets show FedSAP achieving higher mean global accuracy than the strongest baseline while supporting up to 80% client pruning ratios.
By Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao, Qilei Li
The paper introduces a latent information sharing scheme for federated learning that mitigates client drift by sharing a small amount of hidden‑layer activations. The authors demonstrate both theoretically and empirically that this approach improves training efficiency while maintaining convergence guarantees and data privacy. Compared to existing methods such as FedProx, SCAFFOLD, FedPVR, FedProto, and SplitFed, the proposed method achieves higher model accuracy within a fixed round budget without adding significant communication overhead.
By Seungjun Lee, Ensieh Khazaei, Dimitrios Hatzinakos, Baturalp Buyukates, Sunwoo Lee
The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.
By Gianluca Mittone, Marco Aldinucci
arXiv:2608. 14654v1 Announce Type: cross Abstract: Federated Learning (FL) is a collaborative paradigm that enables multiple devices to train a global model while preserving local data privacy.
By Hai Anh Tran, Cuong Ta, Truong X. Tran
arXiv:2607. 13119v1 Announce Type: cross Abstract: In standard federated learning systems, the parameter server broadcasts the global model to the participating devices in every iteration.
By Chung-Hsuan Hu, Zheng Chen, Erik G. Larsson
The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By forming loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to traditional sharded DP approaches.
arXiv:2606. 26822v1 Announce Type: new Abstract: Federated Learning (FL) has become a foundational paradigm for privacy-preserving distributed intelligence, yet its scalability remains fundamentally constrained by communication bottlenecks, device heterogeneity, and the challenges of training under statistically non-IID data.
By Farwa Ikram, Dipanwita Thakur, Antonella Guzzo, Giancarlo Fortino
arXiv:2508. 06692v2 Announce Type: replace Abstract: Federated learning systems typically allocate gradient compression by link speed.
By Md. Akmol Masud, Md Abrar Jahin, Mahmud Hasan
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
By Amir Sarfi, Benjamin Th\'erien, Joel Lidin, Eugene Belilovsky
arXiv:2604. 25421v2 Announce Type: replace-cross Abstract: Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data, yet in mobile deployments the training wall-clock is often bottlenecked by straggler-limited uplink communication under heterogeneous bandwidth and intermittent participation.
By Changyu Li, Shuanghong Huang, Jiashen Liu, Ming Lei, Jidu Xing, Kaishun Wu, Lu Wang, Fei Luo