arXiv Computer Vision

DensityKV: Density-Guided KV Cache Compression for Long Video Generation

DensityKV is a training‑free strategy for managing the historical key‑value (KV) cache in autoregressive video diffusion models. It creates a separate token‑level KV bank for each attention head and uses Soft‑Riesz density to measure and limit local redundancy among post‑RoPE keys, thereby preventing the KV archive from growing indefinitely. Experiments on three video generation backbones demonstrate that, with the same KV capacity limit, DensityKV improves long‑horizon consistency and generation stability while keeping persistent storage bounded regardless of rollout length.

arXiv AI
Jun 12

TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.

By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
arXiv Machine Learning
Sep 24

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

DeltaS is a query‑agnostic, training‑free method for evicting key‑value cache entries in hybrid video‑language models that combine linear and full attention. It uses the change in the recurrent state of gated‑delta linear attention—called state drift—to decide which video chunks to keep, selecting those that induce larger normalized state changes. In experiments with a fixed memory budget, DeltaS outperforms position‑, attention‑, and key‑value‑based eviction signals, improving performance by 2.1 points on average across six long‑video benchmarks and 5.6 points on the longest benchmark, while adding only 1.9% of the forward‑pass cost.

By Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim
Hugging Face Trending Papers
Jun 11

TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift.

arXiv AI
Jun 2

STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models

arXiv:2606. 01790v1 Announce Type: cross Abstract: Vision-language-model-based graphical user interface (GUI) agents have shown broad automation capabilities, yet deployment is bottlenecked by a key-value (KV) cache that grows linearly with interaction steps.

By Yuhang Han, Wenzheng Yang, Yujie Chen, Xiangqi Jin, Yaojie Zhang, Siteng Huang, Linfeng Zhang
arXiv Computer Vision
Aug 31

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.

By Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang
arXiv Computer Vision
Aug 31

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.

By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
arXiv Machine Learning
Jul 23

HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

arXiv:2607. 20125v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high resolution where each frame contributes many tokens.

By Jinliang Shen, Lianghao Su, Zheming Li, Kang He, ZiLiang Lai, Yanbing Jiang, Chengru Song