arXiv:2506.10911v2 Announce Type: replace
Abstract: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect...
By Jari Kolehmainen, Nikolay Blagoev, Semih Kara, John Donaghy, Christopher Nies, O\u{g}uzhan Ersoy
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
By Amir Sarfi, Benjamin Th\'erien, Joel Lidin, Eugene Belilovsky
arXiv:2609.36662v1 Announce Type: cross
Abstract: The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of acc...
By Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang
The paper introduces Local Superior Soups, a model‑interpolation based local training technique designed to improve the adaptation of large pre‑trained models in cross‑silo federated learning. By encouraging exploration of a connected low‑loss basin through regularized interpolation, the method reduces the number of communication rounds needed and boosts performance across several widely used FL datasets. The authors provide code for reproducibility.
By Minghui Chen, Meirui Jiang, Xin Zhang, Qi Dou, Zehua Wang, Xiaoxiao Li
arXiv:2302. 09832v4 Announce Type: replace Abstract: In distributed optimization and federated learning, slow and costly communication between parallel devices and the central server constitutes the primary bottleneck.
By Laurent Condat, Ivan Agarsk\'y, Grigory Malinovsky, Peter Richt\'arik
FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.
By Junkang Liu