arXiv Machine Learning By Jinliang Shen, Lianghao Su, Zheming Li, Kang He, ZiLiang Lai, Yanbing Jiang, Chengru Song

HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

Read the original on arXiv Machine Learning →

arXiv:2607. 20125v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high resolution where each frame contributes many tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
3d ago

In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

arXiv:2609.32540v2 Announce Type: replace-cross Abstract: Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, eac...

By Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang
arXiv Machine Learning
Sep 24

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

DeltaS is a query‑agnostic, training‑free method for evicting key‑value cache entries in hybrid video‑language models that combine linear and full attention. It uses the change in the recurrent state of gated‑delta linear attention—called state drift—to decide which video chunks to keep, selecting those that induce larger normalized state changes. In experiments with a fixed memory budget, DeltaS outperforms position‑, attention‑, and key‑value‑based eviction signals, improving performance by 2.1 points on average across six long‑video benchmarks and 5.6 points on the longest benchmark, while adding only 1.9% of the forward‑pass cost.

By Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim
arXiv Computer Vision
Sep 4

StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs

StreamTTT is a streaming vision-language model that balances real-time perception with long-term memory by writing long-range history into fast weights outside the attention context, while keeping a short sliding key-value cache for recent evidence. The model is trained on both offline long-video QA and a new real-time QA corpus, and it outperforms SimpleStream-4B on OVO-Bench by 1.4 points in real-time perception and 3.7 points in backward tracing. StreamTTT-4B also competes with the larger SimpleStream-8B on the StreamingBench Real-Time Visual Understanding subset.

By Joya Chen, Zeyun Zhong, Mike Zheng Shou