In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compression ratio is often accompanied by the loss of critical details.
arXiv:2608.24293v1 Announce Type: new
Abstract: Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with varia...
By Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim, Donghoon Lee
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.
Token-Budget Distillation (TBD) is a parameter‑efficient fine‑tuning framework that adapts video vision‑language models to a fixed token budget. It freezes the pretrained backbone, updates only LoRA adapters, and incorporates FlashVID visual token compression. TBD uses a dual‑path teacher‑student design with full‑token supervision and compressed student optimization, enabling the student to recover full‑token semantics while remaining efficient under aggressive token reduction.
By Xiaoyang Guo, Guoping Luo, Jusheng Zhang, Keze Wang, Wenhao Wang
arXiv:2606.06158v2 Announce Type: replace
Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous...
By Kevin Dave, Sai Aditya Patkuri, Chhaya Kumar Das, Gouranga Bala, Rajeshkumar SA, R. Venkatesh Babu
arXiv:2603. 00198v2 Announce Type: replace-cross Abstract: Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated as token pruning.
By Jindong Jiang, Amala Sanjay Deshmukh, Kateryna Chumachenko, Karan Sapra, Zhiding Yu, Guilin Liu, Andrew Tao, Pavlo Molchanov, Jan Kautz, Wonmin Byeon
arXiv:2607.14088v2 Announce Type: replace
Abstract: Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure...
By Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang
arXiv:2609.37042v1 Announce Type: cross
Abstract: Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the lar...
By Shuo Yang, Changbai Li, Rui Tang, Xinyu Zhao, Linlin Yang, Baochang Zhang
arXiv:2608.20473v1 Announce Type: new
Abstract: Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is there...
By Wenti Yin, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Changxin Gao, Nong Sang
arXiv:2609.16722v1 Announce Type: new
Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates co...
By Haoyu Guo, Yuan Feng, Junlin Lv, Mingjun Xiao, S Kevin Zhou, Xike Xie
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
arXiv:2610.01785v1 Announce Type: cross
Abstract: Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibit...
By Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu