arXiv Machine Learning By Haobo Zhang, Jiankun Wang, Suraj Rajendran, Weishen Pan, Lam Tsoi, Yong Chen, Fei Wang, Jiayu Zhou

Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated Learning

Read the original on arXiv Machine Learning →

arXiv:2607. 14367v1 Announce Type: new Abstract: Federated fine-tuning of large pre-trained models increasingly relies on Low-Rank Adaptation (LoRA) to reduce communication and computation, but heterogeneous clients can make adapter aggregation unstable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.

By Junkang Liu
arXiv AI
Sep 10

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

FedSubMuon introduces a communication‑efficient federated fine‑tuning approach for large language models by optimizing compact coefficient matrices within shared structured subspaces, thereby keeping Muon’s matrix‑aware optimization while reducing client upload size. An accuracy‑oriented variant, FedSubMuon‑GT, further adapts subspace bases using projected gradients to better align with task‑relevant directions. Experiments on instruction tuning and mathematical reasoning demonstrate that FedSubMuon‑GT achieves the best overall accuracy on most dataset‑model pairs, while FedSubMuon outperforms all matched‑budget baselines and reduces communication by up to 5.5× on Llama‑1B and 1.4× on Qwen‑4B compared to the closest baseline.

By Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang
arXiv Machine Learning
1d ago

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

FedFit introduces a federated fine‑tuning framework for large language models that reduces communication overhead by using a disjoint shared vector‑bank parameterization to reconstruct adapter matrices from two compact global vector banks. It resolves the aggregation dilemma between Sum‑of‑Products and Product‑of‑Sums through an alternating optimization schedule that alternates between accurate single‑bank updates and joint updates corrected by a Residual Spectral Aggregation mechanism. The method also incorporates blockwise quantization with client‑side error feedback and provides theoretical convergence guarantees, achieving perplexity comparable to standard federated LoRA while delivering up to 100× higher compression ratios on Qwen2.5 models.

By Hang Zou, Chao Zhang, Yuzhi Yang, Yu Tian, Samson Lasaulce, M\'erouane Debbah