Block Sparse Matrices for Smaller and Faster Language Models
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
arXiv:2607. 20434v1 Announce Type: cross Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.
arXiv:2606. 02328v1 Announce Type: new Abstract: We explore Riemannian optimization techniques for rank-factored matrix parameters, targeting contemporary deep learning applications.
The paper studies matrices built from block‑diagonal factors interleaved with fixed permutations, a structured family useful in deep learning for balancing expressivity and efficiency. By applying Riemannian geometry, the authors determine when this class forms a smooth manifold and develop Riemannian tools for the orthogonal two‑factor case. They propose efficient algorithms that use automatic differentiation, allow parameter sharing, and avoid dense matrix construction, testing them on matrix approximation and fine‑tuning large language models, while also exploring properties of factorizations with more factors.