arXiv Machine Learning

HeteRo-Select: Informativeness as the Participation Driver in Heterogeneous Federated Learning

arXiv:2508. 06692v2 Announce Type: replace Abstract: Federated learning systems typically allocate gradient compression by link speed.

arXiv Machine Learning
Sep 2

Contribution-Aware Bandwidth Allocation for Multimodal Split Learning

The paper introduces ModalShare, a bandwidth allocation method for multimodal split learning that assigns each modality a keep‑ratio based on its Shapley contribution score. Unlike existing compression schemes that split the uplink budget proportionally to activation size, ModalShare explicitly optimizes the split across modalities, requiring no extra uplink traffic or client computation. Experiments on CREMA‑D and MVSA datasets show that ModalShare improves accuracy by 12.4–15.4 percentage points over equal keep‑ratios under a 5× compression budget, outperforming three compressors across multiple datasets and budgets.

By Iason Ofeidis, Leandros Tassiulas
arXiv Machine Learning
1d ago

FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices

FedSAP is a federated learning framework that addresses heterogeneous edge devices by using structured pruning as a budget-constrained tri-state channel allocation. It partitions model channels into a Global pool, pseudo-domain-specific Private pools, and a Dropped state, allowing broadly useful features to be shared while isolating domain-sensitive updates. Experiments on Digits and Office-Caltech datasets show FedSAP achieving higher mean global accuracy than the strongest baseline while supporting up to 80% client pruning ratios.

By Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao, Qilei Li
arXiv AI
Jul 22

Federated Lightweight Fine-Tuning

arXiv:2607. 18343v1 Announce Type: cross Abstract: Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor.

By Radhakrishna Achanta, Will Reed
arXiv Machine Learning
Sep 14

Hidden in Rounds: Predicting the Time Cost of 802.11 Contention in Federated Learning

The paper investigates how the time cost of 802.11 contention affects federated learning. Using ns-3 simulations, the authors measure frame-delivery ratios and saturation throughput across various client densities and offered loads, then employ a FedAvg trainer that uses these ratios to estimate communication time. Across 720 runs with diverse datasets, partitions, densities, loads, and seeds, all models reached target accuracy within the round budget, with communication time-to-target increasing significantly as client density rose. The study also compares uniform and persistent heterogeneous participation, finding no statistically significant accuracy gap, though confidence intervals are wide. The results are specific to the evaluated configurations and do not generalize to all convergence or fairness scenarios.

By Satwat Bashir, Tasos Dagiuklas
arXiv AI
Sep 18

Accelerating Sharded Data Parallelism at Scale with Federated Learning

The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.

By Gianluca Mittone, Marco Aldinucci
arXiv Machine Learning
Aug 31

Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning

The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.

By Yiming Xie, Lili Su, Ningfang Mi