arXiv Machine Learning By Yijun Lu, Zihan Fang, Pengpeng Qiao, Zheng Lin, Jing Yang, Yuxin Zhang, Por Lip Yee, Zhe Chen, Jun Luo

Conflict-Aware Federated Fine-Tuning of Large Language Models with Mixture-of-Experts

Read the original on arXiv Machine Learning →

arXiv:2606. 15625v1 Announce Type: new Abstract: The continuous scaling of large language models (LLMs) incurs prohibitive computational costs, making Mixture-of-Experts (MoE) a scalable alternative for efficient fine-tuning via sparse activation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 15

Task-Aware Federated Fine-Tuning for MoE-based Large Language Models

The paper introduces FedTAR, a task-aware federated fine‑tuning approach for Mixture‑of‑Experts (MoE) large language models. FedTAR links local client updates to task preferences using routing outputs and Singular Value Decomposition to extract low‑dimensional task coordinates and update directions. It then aggregates updates within and across task clusters, reconstructing the final update to preserve expert specialization and reduce interference, achieving state‑of‑the‑art performance on four benchmark tasks under non‑IID settings.

By Tingqi Wang, Hongyu Ke, Haoxin Wang, Rafal Angryk, Zhipeng Cai
arXiv AI
Sep 2

Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity

The paper introduces FedRoRA, a federated learning framework that combines Low‑Rank Adaptation (LoRA) with rank‑heterogeneous personalization. It separates model adaptation into shared global directions and client‑specific rank‑wise magnitudes, using SVD on the server to extract a global subspace and a personalized projection with top‑k selection for each client. Experiments on natural language understanding and generation tasks show that FedRoRA outperforms existing state‑of‑the‑art methods.

By Lei Wang, Jieming Bian, Letian Zhang, Jie Xu
arXiv Machine Learning
1d ago

Latent Information Sharing for Accelerating Federated Learning

The paper introduces a latent information sharing scheme for federated learning that mitigates client drift by sharing a small amount of hidden‑layer activations. The authors demonstrate both theoretically and empirically that this approach improves training efficiency while maintaining convergence guarantees and data privacy. Compared to existing methods such as FedProx, SCAFFOLD, FedPVR, FedProto, and SplitFed, the proposed method achieves higher model accuracy within a fixed round budget without adding significant communication overhead.

By Seungjun Lee, Ensieh Khazaei, Dimitrios Hatzinakos, Baturalp Buyukates, Sunwoo Lee
arXiv Machine Learning
Aug 18

Global Federated Learning Strategies for Building Efficient Personalized Models

arXiv:2608. 15107v1 Announce Type: new Abstract: Federated learning (FL) is a practical framework that can train models on distributed user data while guaranteeing data privacy; however, due to heterogeneity in which each user has a different data distribution, problems frequently arise where both global and personalization performance deteriorate simultaneously.

By Seongyoon Kim