Overcoming the Communication-Performance Tradeoff in LLM Pretraining
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
arXiv:2607. 01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reconstruction overhead increase with model size and system scale.
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
arXiv:2506.10911v2 Announce Type: replace Abstract: Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect...
arXiv:2609.37899v1 Announce Type: new Abstract: Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient...
arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.
arXiv:2606. 04511v1 Announce Type: cross Abstract: Sparse attention reduces compute and memory bandwidth for long-context LLM inference.
The paper introduces a communication‑efficient method for adapting large language models on decentralized GPU meshes. It proposes an asynchronous two‑circuit system that uses fast compressed training with activation masking for pipeline‑parallel transfer and compressed data‑parallel synchronization, while a slower anchor circuit performs occasional unmasked passes. A spectral correction optimizer then denoises the masked gradients using these anchor priors, enabling high compression rates and achieving up to 40× throughput gains over internet‑grade connections while matching dense uncompressed performance.
The paper introduces the "Compression Trinity," a unified framework that jointly applies sparsity, quantization, and low‑rank approximations to compress large language models. It presents several methods—MKOR, SLoPe, OPTIMA, PATCH, and SLiM—that leverage these three pillars to accelerate training, reduce memory bandwidth, and recover accuracy, achieving significant speedups and accuracy gains over existing techniques. The results demonstrate that combining all three compression strategies is essential for efficient, scalable, high‑performance LLM deployment.
Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) a...
arXiv:2602. 04396v2 Announce Type: replace-cross Abstract: Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth.
arXiv:2502. 11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining.
The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.
arXiv:2609.06557v1 Announce Type: new Abstract: Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires stic...