Block Sparse Matrices for Smaller and Faster Language Models
Related stories
Efficient training of language models to fill in the middle
Large Language Models: A New Moore's Law?
Break Through the Compression Bottleneck: From Theory to Practice
arXiv:2607. 20434v1 Announce Type: cross Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.
Riemannian Gradient Descent for Low-Rank Architectures
arXiv:2606. 02328v1 Announce Type: new Abstract: We explore Riemannian optimization techniques for rank-factored matrix parameters, targeting contemporary deep learning applications.
Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices
The paper studies matrices built from block‑diagonal factors interleaved with fixed permutations, a structured family useful in deep learning for balancing expressivity and efficiency. By applying Riemannian geometry, the authors determine when this class forms a smooth manifold and develop Riemannian tools for the orthogonal two‑factor case. They propose efficient algorithms that use automatic differentiation, allow parameter sharing, and avoid dense matrix construction, testing them on matrix approximation and fine‑tuning large language models, while also exploring properties of factorizations with more factors.
Evaluating large language models trained on code
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
Databricks ❤️ Hugging Face: up to 40% faster training and tuning of Large Language Models
Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs
Residual sparsification via output importance (PARSER) is a new compression technique for mixture-of-experts large language models that shifts the compression objective from minimizing isolated matrix errors to preserving the expert output error. By introducing output importance, PARSER measures each residual’s contribution to the final expert output and compresses accordingly. Experiments show that PARSER reduces the accuracy gap to the uncompressed model by 1.41× on Qwen and 1.44× on DeepSeek while achieving the same peak memory reduction.
Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models
arXiv:2609.12475v1 Announce Type: new Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluati...
