arXiv Computer Vision By Yangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song

Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation

Read the original on arXiv Computer Vision →

The paper introduces AnchorCache, a parameter‑free token‑layout and attention‑mask design that decouples reference tokens from the target in in‑context diffusion transformers. By inserting static text anchors, the method conditions reference representations on the instruction during cache construction, enabling exact key‑value reuse across denoising steps. To restore quality lost by this structural change, the authors employ teacher‑forced velocity distillation followed by a brief on‑policy stage, achieving full‑attention quality while delivering up to 6.40× speedup in diffusion transformer inference across image, speech, and video benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv Computer Vision
6d ago

Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion

The paper introduces HetA-DiT, a heterogeneous attention mechanism for video diffusion models that allocates computation based on token difficulty. A lightweight uncertainty branch predicts denoising difficulty, routing uncertain tokens through dense global attention while applying efficient local attention to reliable tokens. This adaptive routing retains global context where needed, offers a single parameter to balance quality and efficiency, and achieves competitive generation quality while only about 20% of tokens use dense attention.

By Olga Zatsarynna, Denis Korzhenkov, Juergen Gall, Amir Habibian, Mohsen Ghafoorian
arXiv AI
Jul 8

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

arXiv:2607. 06523v1 Announce Type: new Abstract: Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation.

By Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jesus Olivera
arXiv Computer Vision
3d ago

Looped Diffusion Transformer

arXiv:2609.40305v1 Announce Type: new Abstract: Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternat...

By Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang