FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.
By Junkang Liu
arXiv:2506.10911v2 Announce Type: replace
Abstract: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect...
By Jari Kolehmainen, Nikolay Blagoev, Semih Kara, John Donaghy, Christopher Nies, O\u{g}uzhan Ersoy
arXiv:2607. 01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reconstruction overhead increase with model size and system scale.
By Mingkai Zheng, Junlin Chen, Haotian Xie, Zhao Zhang
arXiv:2604. 24012v3 Announce Type: replace Abstract: Federated learning enables a population of clients to collaboratively train machine learning models without exchanging their raw data, but standard algorithms such as FedAvg suffer from slow convergence and high communication and memory costs in heterogeneous, resource-constrained environments.
By Yutong He, Zhengyang Huang, Jiahe Geng, Kun Yuan
The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.
By Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini
arXiv:2608. 06563v1 Announce Type: new Abstract: Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications.
By Grigory Malinovsky
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
By Amir Sarfi, Benjamin Th\'erien, Joel Lidin, Eugene Belilovsky
arXiv:2405. 11667v2 Announce Type: replace Abstract: Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD.
By Kumar Kshitij Patel, Margalit Glasgow, Ali Zindari, Lingxiao Wang, Sebastian U. Stich, Ziheng Cheng, Nirmit Joshi, Nathan Srebro
arXiv:2601. 16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments.
By Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu
arXiv:2609.36301v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regim...
By Honam Wong, Surbhi Goel, Enric Boix-Adser\`a
arXiv:2606. 16384v1 Announce Type: new Abstract: Pretraining language models with extended context windows enhances their ability to leverage rich information during generation.
By Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi, Gil Avraham, Violetta Shevchenko, Yan Zuo, Chamin Hewa Koneputugodage, Alexander Long
arXiv:2606. 17526v1 Announce Type: new Abstract: Efficient optimization is essential for training large language models.
By Da Chang, Ganzhao Yuan