A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
Related stories
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
On Generalisation Error Bounds for Transformers
arXiv:2410. 11500v2 Announce Type: replace-cross Abstract: In this paper, we establish a collection of covering number bounds for linear function classes under various norm constraints on the inputs and matrices.
Memory-efficient Diffusion Transformers with Quanto and Diffusers
On the Expressive Power of Transformers
arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.
Modern analog computing for solving differential and matrix equations
arXiv:2606. 13179v1 Announce Type: cross Abstract: In recent years, driven by the computational demands of data-intensive applications such as artificial intelligence and scientific computing, analog computing has gained renewed interest.
Native-speed vLLM transformers modeling backend
Modern analog computing for solving differential and matrix equations
In recent years, driven by the computational demands of data-intensive applications such as artificial intelligence and scientific computing, analog computing has gained renewed interest. Given the diversity of computational tasks and recent advancements in analog CMOS circuits and resistive memory technologies, we refer to the evolving landscape as modern analog computing.
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers
arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.
Transforms for LLM Quantization: The Great Inversion and Format Co-Design
The paper surveys the use of linear, functionâpreserving transforms in 4âbit largeâlanguageâmodel (LLM) quantization, formalizing the underlying principle as the "Great Inversion"âthe tradeâoff between energy concentration favored by allocationâflexible coding and withinâgroup flattening favored by grouped sharedâscale quantization. It reviews 200 works, classifies 43 transform methods by structure, dataâawareness, construction approach, and runtime cost, and examines how they interact with GPTQ rounding. The study also explores how different number formats (FP4, MXFP4, NVFP4) influence the optimal transform choice and outlines open research problems. "whyItMatters":"The survey clarifies the conflicting objectives in transformâbased LLM quantization and provides a practical guide for selecting transforms based on deployment regime, thereby informing future research and deployment strategies."
Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition
arXiv:2411. 09816v5 Announce Type: replace Abstract: Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices.
Transformer Heads Looking for Order
The paper demonstrates that a single-head, single-layer transformer cannot determine whether a bit sequence is ordered, whereas a two-head, single-layer transformer can. This distinction is shown under a model where transformers include an output MLP. The study provides a concrete example of how increasing the number of heads can enhance a transformerâs computational capability.