Accelerated Inference with Optimum and Transformers Pipelines
Related stories
Sparse Layers are Critical to Scaling Looped Language Models
arXiv:2605. 09165v2 Announce Type: replace Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries.
Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layersâ role in information extraction and characterizes the tradeâoff between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
Recirculation
The paper introduces recirculation, an inferenceâtime architectural enhancement for foundation models that reduces perplexity and improves accuracy on generation and reasoning tasks without adding significant latency. Recirculation adds a specific form of recurrence, enabling the model to function as a dynamical system that tracks belief states, and is distinct from chainâofâthought or depthârecurrence methods. An adaptive variant requires minimal hyperparameter tuning and achieves notable gains on the Gemma3 family, including a 23% perplexity drop and a 21% accuracy increase on GSM8k.
Fixed Universal Transformers
The paper introduces fixed universal transformers, which are transformers with immutable internal parameters that can emulate any transformer within a specified class by encoding the target modelâs description into the input embedding. The authors provide explicit sparse constructions that achieve universality when the embedding dimension is large enough, and demonstrate that universality is genericârandomly initialized transformers are almost surely universal. Empirical tests on parenthesis balancing and multiâhop reasoning tasks support the theory, suggesting that a transformerâs expressive power largely stems from its input representation rather than its learned weights.
FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates
FlashLoop is a trainingâfree inference framework for Looped Transformers that reduces crossâloop redundancy by employing tokenâsparse updates, sparse attention, and KVâresidual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to lowâbit quantization. The method achieves lossless accuracy with up to 1.64Ă speedup and 6Ă KVâcache memory reduction across several Looped Transformer models.
ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
arXiv:2607. 23478v1 Announce Type: cross Abstract: Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under FHE remains prohibitively expensive.
Up to 3.2x Faster Inference with LFM2.5-DSpark
Amortized quadrature for posterior expectations in inverse problems
arXiv:2606.15871v2 Announce Type: replace-cross Abstract: Uncertainty in the solution of an inverse problem and in the tasks performed on it is quantified by posterior expectations, each an average o...
Scaling up BERT-like model Inference on modern CPU - Part 2
FPTQuant: Function-Preserving Transforms for LLM Quantization
arXiv:2506. 04985v2 Announce Type: replace Abstract: Large language models (LLMs) require substantial compute, and thus energy, at inference time.
Large-scale empirical tuning and comparison of default optimizers for variational inference
arXiv:2606. 07841v1 Announce Type: cross Abstract: Black-box variational inference (BBVI) is a methodology for posterior approximation that relies on stochastic optimization.