Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
Related stories
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Accelerating PyTorch Transformers with Intel Sapphire Rapids - part 2
Accelerating PyTorch Transformers with Intel Sapphire Rapids - part 1
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
DC-Gen is a post‑training framework that accelerates text‑to‑image diffusion models by using a deeply compressed latent space. It first aligns the base model’s latent representations with a lightweight embedding alignment, then applies minimal LoRA fine‑tuning to preserve generation quality. Experiments on SANA and FLUX.1‑Krea show that DC‑Gen‑FLUX cuts 4K image generation latency by 53× on an NVIDIA H100 and, with NVFP4 SVDQuant, achieves a 138× total speedup on a single NVIDIA 5090 GPU.
Native-speed vLLM transformers modeling backend
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
arXiv:2608. 19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets.
Accelerating Hugging Face Transformers with AWS Inferentia2
Mistral AI partners with NVIDIA to accelerate open frontier models
Fine-Tuning of Transformer models with Frames
The paper introduces FrameFT, a parameter-efficient fine-tuning method for transformer models that represents weight updates using sparse coefficients in a Fusion Frame basis. This approach reduces memory usage by storing only the sparse coefficients, leading to significant compute advantages and formal convergence guarantees. Experiments on language and vision tasks show that FrameFT matches or surpasses state‑of‑the‑art PEFT techniques while requiring far fewer trainable parameters.
Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey
arXiv:2609.01212v1 Announce Type: new Abstract: With the rapid and continuous growth in the incorporation of machine learning models based on the Transformer architecture, capable deployment is in hi...