arXiv:2505. 10946v3 Announce Type: replace-cross Abstract: Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation units across modalities.
By Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Robert Schober, Deniz G\"und\"uz
arXiv:2607. 29363v1 Announce Type: cross Abstract: Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation.
By Yi Luo, Rongzhi Gu, Jixun Yao
The paper introduces TokCode, a token encoding framework that enhances robustness in generative semantic communication by restructuring redundancy in the semantic domain. TokCode leverages a lightweight adapter to transform a large language model into a token encoder, avoiding the need for a dedicated deep model. A channel-quality-aware distillation method (CADET) trains the adapter across diverse erasure rates, producing a reconfigurable low‑rank adapter that enables efficient reinforcement learning and achieves significant improvements in image similarity over existing receiver‑side recovery benchmarks.
By Jingzhi Hu, Ouya Wang, Geoffrey Ye Li
arXiv:2602. 12338v2 Announce Type: replace Abstract: Token Communications (TokenCom) has recently emerged as an effective new paradigm, where tokens are the unified units of multimodal communications and computations, enabling efficient digital semantic- and goal-oriented communications in future wireless networks.
By Farshad Zeinali, Mahdi Boloursaz Mashhadi, Rahim Tafazolli
arXiv:2606.06158v2 Announce Type: replace
Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous...
By Kevin Dave, Sai Aditya Patkuri, Chhaya Kumar Das, Gouranga Bala, Rajeshkumar SA, R. Venkatesh Babu
The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.
By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim