arXiv AI

Region-Adaptive Sampling for Diffusion Transformers

arXiv:2502. 10389v2 Announce Type: replace-cross Abstract: Diffusion models (DMs) have become the leading choice for generative tasks across diverse domains.

arXiv AI
Jun 15

HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing

arXiv:2606. 13898v1 Announce Type: cross Abstract: Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer use and account for a major share of traffic in Photoshop and Lightroom.

By Haoran You, Yotam Nitzan, Lingzhi Zhang, Yifan Gong, Mang-Tik Chiu, Connelly Barnes, Yan Kang, Yuqian Zhou, Eli Shechtman, Sohrab Amirghodsi
arXiv AI
Aug 25

ChebBooster: A Training-Free Approach for Efficient Diffusion Transformer Inference via Chebyshev-Inspired Extrapolation

ChebBooster is a training‑free extrapolation framework that accelerates Diffusion Transformers (DiTs) by using Chebyshev polynomial theory. It employs a Barycentric formulation for numerically stable evaluation and separates the process into an offline weight precomputation phase and a lightweight online application stage. Experiments on DiT‑XL/2, PixArt‑Σ, and FLUX.1‑dev show consistent visual quality gains and up to 3.68× latency speedup and 5.12× FLOPs reduction compared to existing training‑free baselines.

By Chengjie Lu, Tianchi Deng, Zhengqi He, Chengwen Luo, Xueliang Li
arXiv AI
Jul 13

Transition Matching Distillation for Fast Video Generation

arXiv:2601. 09881v2 Announce Type: replace-cross Abstract: Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process.

By Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, Arash Vahdat
arXiv Computer Vision
5d ago

Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion

The paper introduces HetA-DiT, a heterogeneous attention mechanism for video diffusion models that allocates computation based on token difficulty. A lightweight uncertainty branch predicts denoising difficulty, routing uncertain tokens through dense global attention while applying efficient local attention to reliable tokens. This adaptive routing retains global context where needed, offers a single parameter to balance quality and efficiency, and achieves competitive generation quality while only about 20% of tokens use dense attention.

By Olga Zatsarynna, Denis Korzhenkov, Juergen Gall, Amir Habibian, Mohsen Ghafoorian
arXiv Computer Vision
2d ago

Looped Diffusion Transformer

arXiv:2609.40305v1 Announce Type: new Abstract: Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternat...

By Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang