Hugging Face Trending Papers

Unifying Local Communications and Local Updates for LLM Pretraining

Read the original on Hugging Face Trending Papers →

Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data centers, and lower-bandwidth links. Many practical methods reduce communication frequency but still rely on synchronous All-Reduce operations that maintain identical model states and tie progress to global collectives.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
1d ago

SeedFlood: A Step Toward Scalable Decentralized Fine-Tuning of LLMs

SeedFlood is a novel decentralized fine‑tuning method for large language models that scales to billions of parameters and hundreds of clients. It leverages the seed‑reconstructible structure of zeroth‑order gradients to reduce message sizes to near‑zero, enabling efficient flooding across the network. Experiments show SeedFlood outperforms standard zeroth‑order baselines in communication efficiency and generalization, and rivals first‑order gossip methods while incurring far less communication cost.

By Jihun Kim, Dongyeop Lee, Namhoon Lee
arXiv Machine Learning
Jun 16

Photon: Federated LLM Pre-Training

arXiv:2411. 02908v2 Announce Type: replace Abstract: Scaling large language models (LLMs) demands extensive data and computing resources, which are traditionally constrained to data centers by the high-bandwidth requirements of distributed training.

By Lorenzo Sani, Alex Iacob, Zeyu Cao, Royson Lee, Bill Marino, Yan Gao, Dongqi Cai, Zexi Li, Wanru Zhao, Xinchi Qiu, Nicholas D. Lane