Accelerating PyTorch Transformers with Intel Sapphire Rapids - part 2
Related stories
Accelerating PyTorch distributed fine-tuning with Intel technologies
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey
arXiv:2609.01212v1 Announce Type: new Abstract: With the rapid and continuous growth in the incorporation of machine learning models based on the Transformer architecture, capable deployment is in hi...
Hugging Face and Graphcore partner for IPU-optimized Transformers
FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers
arXiv:2606. 31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers.
How 🤗 Accelerate runs very large models thanks to PyTorch
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers
ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator
ShatterQuant is a hardware-software co-designed framework that enables mixed-precision quantization within individual tensors by assigning different bit-widths to blocks of a weight projection. It couples precision granularity with processing element configuration, allowing each precision to determine an effective block height. The framework includes a hardware-aware post-training method based on block-level standard deviation and weight sensitivity, a ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities, and an evaluation showing 1.5 TOPS, 760 GOPS/$mm^2$ area efficiency, and 2.8 TOPS/W energy efficiency on a TSMC 16nm PDK implementation.
How we sped up transformer inference 100x for 🤗 API customers
Tile-Level Activation Overlap for Efficient LLM Inference
arXiv:2607. 02521v1 Announce Type: cross Abstract: SwiGLU is the dominant MLP activation in modern large language models, yet its intermediate tensor materialization costs 9-37% of MLP execution time.