Accelerate Large Model Training using DeepSpeed
Related stories
Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
arXiv:2609.21672v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeab...
Training Variable Long Sequences with Data-Centric Parallel
arXiv:2608. 07524v1 Announce Type: new Abstract: Training deep learning models on variable long sequences poses significant computational challenges.
Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models
The paper introduces Recursive Block-Diagonal Coupling (RBDC), a training protocol that builds wide vision models by recursively coupling narrower, independently trained models in a parameter‑free block‑diagonal manner. RBDC allows flexible allocation of training budgets across all models and, when applied to vision transformers (DeiT) and convolutional networks (ResNet) on ImageNet, achieves a 30% reduction in FLOPs while maintaining similar test accuracies. Additionally, models trained with RBDC outperform those from existing growth methods at the same training FLOPs and serve as stronger backbones for downstream tasks such as object detection and instance segmentation.
GaLore: Advancing Large Model Training on Consumer-grade Hardware
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
arXiv:2606. 00888v1 Announce Type: cross Abstract: Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates.
Efficient Scaling of LLM Training with Flexible Context Parallelism
arXiv:2602. 21788v2 Announce Type: replace-cross Abstract: Scaling long-context capabilities is crucial for Large Language Models (LLMs).
Scaling depth capacity via zero/one-layer model expansion
arXiv:2511. 04981v2 Announce Type: replace Abstract: Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost.
Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks
Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developments, training and deployment of neural network models on embedding and edge devices face significant challenges due to limited memory and computational resources.
$\mu$pscaling small models: Principled warm starts and hyperparameter transfer
arXiv:2602. 10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets.
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
arXiv:2402.11215v4 Announce Type: replace Abstract: The choice of batch size in minibatch stochastic gradient optimization is critical for both optimization and generalization performance in large-sc...