arXiv:2608. 16192v1 Announce Type: new Abstract: Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver.
By Jia Guo, Xiaohan Zhao, Changwang Liu, Shuqing He, Chenyang Zhang, Bingchuan Zhao, Jinqi Zhu
arXiv:2608. 09176v1 Announce Type: cross Abstract: Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost.
By Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang
The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.
By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim
arXiv:2608. 15694v1 Announce Type: cross Abstract: Conditional image-to-image generators are single-shot: they map input features to an output in one forward pass and treat it as final, with no opportunity to improve on it.
By Kareem Hassani, Chaymaa Abbas, Hadi Al Mubasher, Mariette Awad
The paper introduces TokCode, a token encoding framework that enhances robustness in generative semantic communication by restructuring redundancy in the semantic domain. TokCode leverages a lightweight adapter to transform a large language model into a token encoder, avoiding the need for a dedicated deep model. A channel-quality-aware distillation method (CADET) trains the adapter across diverse erasure rates, producing a reconfigurable low‑rank adapter that enables efficient reinforcement learning and achieves significant improvements in image similarity over existing receiver‑side recovery benchmarks.
By Jingzhi Hu, Ouya Wang, Geoffrey Ye Li
arXiv:2607. 25527v1 Announce Type: cross Abstract: Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities.
By Weiming Zhuang, Jiabo Huang, Jingtao Li, Zhizhong Li, Chen Chen, Sina Sajadmanesh, Lingjuan Lyu