Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
Related stories
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Accelerating PyTorch Transformers with Intel Sapphire Rapids - part 2
Accelerating PyTorch Transformers with Intel Sapphire Rapids - part 1
Native-speed vLLM transformers modeling backend
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
arXiv:2608. 19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets.
Accelerating Hugging Face Transformers with AWS Inferentia2
Mistral AI partners with NVIDIA to accelerate open frontier models
MegaFold: Efficient Training of Next-Generation 3D Attention Protein Models on Cross-Platform GPUs
arXiv:2506. 20686v2 Announce Type: replace-cross Abstract: Recent advances in biomolecular modeling have been catalyzed by models such as AlphaFold3 (AF3), which introduce science-informed changes to the transformer architecture.
Accelerate your models with 🤗 Optimum Intel and OpenVINO
AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers
arXiv:2602. 13357v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure.