Block Sparse Matrices for Smaller and Faster Language Models
Related stories
Efficient training of language models to fill in the middle
Large Language Models: A New Moore's Law?
Break Through the Compression Bottleneck: From Theory to Practice
arXiv:2607. 20434v1 Announce Type: cross Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.
Riemannian Gradient Descent for Low-Rank Architectures
arXiv:2606. 02328v1 Announce Type: new Abstract: We explore Riemannian optimization techniques for rank-factored matrix parameters, targeting contemporary deep learning applications.
Evaluating large language models trained on code
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
Databricks ❤️ Hugging Face: up to 40% faster training and tuning of Large Language Models
The Reformer - Pushing the limits of language modeling
Adaptive Block Diffusion: Resolving Training-Inference Mismatch in Diffusion Language Models
arXiv:2606. 29275v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) are typically trained under fixed context structures, restricting denoising to predetermined token subsets.
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning
arXiv:2608. 05104v1 Announce Type: new Abstract: Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges.
Attention-Based Sampler for Diffusion Language Models
arXiv:2604. 08564v2 Announce Type: replace-cross Abstract: Auto-regressive models (ARMs) have established a dominant paradigm in language modeling.
