The paper introduces a semantic‑aware multi‑level neural video codec designed for low‑latency, task‑oriented video transmission over unreliable channels. It builds on the real‑time DCVC‑RT codec by partitioning encoded representations into packets of varying semantic and feature importance, assigning them to priority streams, and employing an error‑resilient entropy model that removes inter‑packet dependencies. Experiments demonstrate that this framework improves robustness against packet erasures, achieving graceful degradation in less important regions while preserving task‑relevant visual content.
By Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du, Harish Viswanathan
arXiv:2608. 16192v1 Announce Type: new Abstract: Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver.
By Jia Guo, Xiaohan Zhao, Changwang Liu, Shuqing He, Chenyang Zhang, Bingchuan Zhao, Jinqi Zhu
arXiv:2602. 12338v2 Announce Type: replace Abstract: Token Communications (TokenCom) has recently emerged as an effective new paradigm, where tokens are the unified units of multimodal communications and computations, enabling efficient digital semantic- and goal-oriented communications in future wireless networks.
By Farshad Zeinali, Mahdi Boloursaz Mashhadi, Rahim Tafazolli
arXiv:2603. 00198v2 Announce Type: replace-cross Abstract: Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated as token pruning.
By Jindong Jiang, Amala Sanjay Deshmukh, Kateryna Chumachenko, Karan Sapra, Zhiding Yu, Guilin Liu, Andrew Tao, Pavlo Molchanov, Jan Kautz, Wonmin Byeon
Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss, a common challenge in satellite and emergency communications.
The paper introduces Visual Token Coding (VTC), a token compression method for video multimodal large language models that mimics classical video coding by predicting I/P frames and measuring residuals to reduce token redundancy. VTC is extended with dynamic features—Dynamic Resolution Input, Dynamic Token Allocation, and Spatial Coverage Top‑K—forming VTC_Dy, which can be applied to existing MLLMs without additional tuning. Experiments on three MLLMs and multiple video benchmarks show that VTC_Dy retains over 100% of average performance with a 50% token budget and 97.8% with a 25% budget, while the code is publicly available.
By Chenxin Fang, Tao Chen, JunChao You, Jun Peng, Yiyi Zhou, Rongrong Ji
arXiv:2606.06158v2 Announce Type: replace
Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous...
By Kevin Dave, Sai Aditya Patkuri, Chhaya Kumar Das, Gouranga Bala, Rajeshkumar SA, R. Venkatesh Babu
The paper introduces PRESLEY, an end‑to‑end video streaming pipeline that uses generative AI layers to selectively degrade and restore less important regions of a frame. By replacing destructive block removal with adaptive in‑place degradation and signaling block strength via a side channel, PRESLEY achieves significant bitrate savings and improved background quality compared to its predecessor and pristine baselines. The authors also analyze the theoretical headroom of this architecture, quantifying remaining cost‑axis headroom and modeling post‑restoration damage to guide future rate‑distortion‑restoration selection.
By Emanuele Artioli, Farzad Tashtarian, Christian Timmerer
arXiv:2608.24293v1 Announce Type: new
Abstract: Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with varia...
By Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim, Donghoon Lee
arXiv:2609.39296v1 Announce Type: new
Abstract: Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most exist...
By Xiangben Zhu, Caili Guo, Yang Yang, Chuanhong Liu, Meiyi Zhu
arXiv:2505. 10946v3 Announce Type: replace-cross Abstract: Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation units across modalities.
By Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Robert Schober, Deniz G\"und\"uz
COVER is a new video watermarking method that targets codec compression as its primary design goal. It embeds the watermark payload in the latent space of a frozen generative video autoencoder and recovers it by re‑encoding the received video into the same latent space. Using a differentiable codec surrogate bank, COVER achieves high bit accuracy across multiple codecs while keeping marked videos visually close to the originals.
By Yuxin Cao, Hao Yang, Ziqi Ding, Jie Hao, Wei Song