GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining
arXiv:2607. 07494v1 Announce Type: cross Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining.
arXiv:2607. 07494v1 Announce Type: cross Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining.
G$^2$PTQ is a post‑training quantization framework that improves large language models by combining first‑ and second‑order information in a globally supervised, block‑wise optimization. It refreshes gradient and Hessian estimates before each Transformer block and uses a trust‑region scaling mechanism to stabilize gradient steps, preventing exploding weight updates. The method achieves better alignment with full‑precision models and outperforms state‑of‑the‑art baselines across various model families and bit‑widths.
arXiv:2607. 01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reconstruction overhead increase with model size and system scale.
arXiv:2508. 15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.
arXiv:2505. 01043v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved impressive performance across various domains.
arXiv:2601. 22813v2 Announce Type: replace Abstract: The NVFP4 lower-precision format, supported in hardware by NVIDIA Blackwell GPUs, promises to allow, for the first time, end-to-end fully-quantized pre-training of massive models such as LLMs.
arXiv:2607. 24953v1 Announce Type: cross Abstract: Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization.
arXiv:2607. 03763v1 Announce Type: cross Abstract: Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD-based training.
arXiv:2506. 01260v2 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks.
RGSQ introduces a Riemannian geometry‑aware post‑training quantization method for large vision‑language models, treating quantization as a reconstruction problem under a Fisher‑Riemannian metric. It identifies modality‑specific sensitive directions via manifold mappings and applies geometry‑aligned rotations and whitening to steer low‑bit perturbations toward loss‑insensitive axes. Experiments on diverse VLM benchmarks show RGSQ delivers the best accuracy and stability in extremely low‑bit settings, outperforming existing VLM‑aware baselines by up to 5.9% and single‑modality methods by up to 8.6%.
arXiv:2605. 09825v4 Announce Type: replace-cross Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable?
The paper introduces the "Compression Trinity," a unified framework that jointly applies sparsity, quantization, and low‑rank approximations to compress large language models. It presents several methods—MKOR, SLoPe, OPTIMA, PATCH, and SLiM—that leverage these three pillars to accelerate training, reduce memory bandwidth, and recover accuracy, achieving significant speedups and accuracy gains over existing techniques. The results demonstrate that combining all three compression strategies is essential for efficient, scalable, high‑performance LLM deployment.