Accelerating PyTorch Transformers with Intel Sapphire Rapids - part 1
Related stories
Accelerating PyTorch distributed fine-tuning with Intel technologies
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
Hugging Face and Graphcore partner for IPU-optimized Transformers
FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers
arXiv:2606. 31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers.
How 🤗 Accelerate runs very large models thanks to PyTorch
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers
How we sped up transformer inference 100x for 🤗 API customers
Tile-Level Activation Overlap for Efficient LLM Inference
arXiv:2607. 02521v1 Announce Type: cross Abstract: SwiGLU is the dominant MLP activation in modern large language models, yet its intermediate tensor materialization costs 9-37% of MLP execution time.
Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles
arXiv:2607. 10942v1 Announce Type: cross Abstract: Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints.
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch
arXiv:2601. 13994v3 Announce Type: replace-cross Abstract: Differentiable sparse linear algebra is foundational for scientific machine learning, yet PyTorch lacks a unified library for it: torch.