Hugging Face Trending Papers

Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

Read the original on Hugging Face Trending Papers →

Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis. However, these models are computationally expensive in runtime, and their cost scales severely with the number of input references.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
Sep 18

Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

The paper investigates how attention sparsity behaves in autoregressive image generation, finding a distinct diagonal sparsity pattern due to spatial locality of visual tokens. It introduces a diagonal‑aware sparse attention mechanism that skips KV entries along the diagonal within a recent window, achieving up to 3.1× higher throughput and 1.19× lower latency with less than 2% quality loss compared to dense inference.

By Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park
arXiv Computer Vision
Sep 7

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

RefDiT is a new framework for reference-guided image generation that addresses the shortcomings of previous methods when handling complex scenes with multiple objects. It introduces local region guidance by decomposing a single identifier token into attribute-level signals, allowing the model to learn correspondences between tokens and specific regions of a reference image. The approach incorporates a low-rank adapter (LoRA) within a diffusion transformer (DiT) to adjust the inference prompt based on user-provided guidance context, thereby enabling more precise local attribute control.

By Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam