arXiv Computer Vision By Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin, Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics

Read the original on arXiv Computer Vision →

The paper introduces a lightweight model that learns controlled latent dynamics to accelerate masked image generation models (MIGMs). By incorporating previous features and sampled tokens, it regresses the average velocity field of feature evolution, reducing redundancy from bi-directional attention. Applied to Lumina-DiMOO, the method achieves over 4× faster text-to-image generation while preserving quality, advancing the efficiency frontier for MIGMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 18

Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

The paper investigates how attention sparsity behaves in autoregressive image generation, finding a distinct diagonal sparsity pattern due to spatial locality of visual tokens. It introduces a diagonal‑aware sparse attention mechanism that skips KV entries along the diagonal within a recent window, achieving up to 3.1× higher throughput and 1.19× lower latency with less than 2% quality loss compared to dense inference.

By Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park