arXiv Machine Learning By Bingyan Xie, Yongjeong Oh, Zihan Chen, Jihong Park, Yongpeng Wu, Wenjun Zhang

Loss-Resilient Wireless Video Token Communication over Block Fading Channels

Read the original on arXiv Machine Learning →

arXiv:2608. 08698v1 Announce Type: new Abstract: Video token communication represents video content as discrete tokens that differ in their importance to reconstruction and exhibit temporal dependencies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 16

Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission

The paper introduces a semantic‑aware multi‑level neural video codec designed for low‑latency, task‑oriented video transmission over unreliable channels. It builds on the real‑time DCVC‑RT codec by partitioning encoded representations into packets of varying semantic and feature importance, assigning them to priority streams, and employing an error‑resilient entropy model that removes inter‑packet dependencies. Experiments demonstrate that this framework improves robustness against packet erasures, achieving graceful degradation in less important regions while preserving task‑relevant visual content.

By Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du, Harish Viswanathan
arXiv Machine Learning
Jul 24

Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications

arXiv:2602. 12338v2 Announce Type: replace Abstract: Token Communications (TokenCom) has recently emerged as an effective new paradigm, where tokens are the unified units of multimodal communications and computations, enabling efficient digital semantic- and goal-oriented communications in future wireless networks.

By Farshad Zeinali, Mahdi Boloursaz Mashhadi, Rahim Tafazolli
arXiv Computer Vision
Aug 31

Visual Token Coding for Video Multimodal Large Language Models

The paper introduces Visual Token Coding (VTC), a token compression method for video multimodal large language models that mimics classical video coding by predicting I/P frames and measuring residuals to reduce token redundancy. VTC is extended with dynamic features—Dynamic Resolution Input, Dynamic Token Allocation, and Spatial Coverage Top‑K—forming VTC_Dy, which can be applied to existing MLLMs without additional tuning. Experiments on three MLLMs and multiple video benchmarks show that VTC_Dy retains over 100% of average performance with a 50% token budget and 97.8% with a 25% budget, while the code is publicly available.

By Chenxin Fang, Tao Chen, JunChao You, Jun Peng, Yiyi Zhou, Rongrong Ji