arXiv AI By Stefan Horoi, Benjamin Th\'erien, Guy Wolf, Eugene Belilovsky

Can Model Merging Improve Aggregation in DiLoCo?

Read the original on arXiv AI →

arXiv:2607. 03011v1 Announce Type: cross Abstract: Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of significant interest in recent years, with a broad array of methods having been proposed to tackle this problem.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Local Superior Soups: A Catalyst for Model Merging in Cross-Silo Federated Learning

The paper introduces Local Superior Soups, a model‑interpolation based local training technique designed to improve the adaptation of large pre‑trained models in cross‑silo federated learning. By encouraging exploration of a connected low‑loss basin through regularized interpolation, the method reduces the number of communication rounds needed and boosts performance across several widely used FL datasets. The authors provide code for reproducibility.

By Minghui Chen, Meirui Jiang, Xin Zhang, Qi Dou, Zehua Wang, Xiaoxiao Li
arXiv AI
2d ago

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.

By Junkang Liu