The groundbreaking development of generative artificial intelligence (AI) is rapidly boosting the ability to generate content such as images and videos, reshaping communication paradigms. This article introduces generative communications (GenCom), a novel paradigm for 6G networks in which large AI models (LAMs) drive semantic understanding, reasoning, and content generation, embedding these into the communication process.
arXiv:2502.08221v2 Announce Type: replace
Abstract: The growing demand for efficient semantic communication systems capable of managing diverse tasks and adapting to fluctuating channel conditions ha...
By Xiang Chen, Shuying Gan, Chenyuan Feng, Xijun Wang, Tony Q. S. Quek
arXiv:2609.39296v1 Announce Type: new
Abstract: Video semantic communication has attracted increasing attention as a promising approach to improving video transmission efficiency. However, most exist...
By Xiangben Zhu, Caili Guo, Yang Yang, Chuanhong Liu, Meiyi Zhu
arXiv:2607. 09183v1 Announce Type: cross Abstract: The groundbreaking development of generative artificial intelligence (AI) is rapidly boosting the ability to generate content such as images and videos, reshaping communication paradigms.
By Wenjun Zhang, Zhiyong Chen, Tong Wu, Guo Lu, Li Song, Feng Yang, Meixia Tao
GenStream is a semantic streaming framework that replaces dense video frames with compact metadata—skeletal keypoints, camera parameters, and a static 3D background model—to enable generative reconstruction of human figures on the client side. By transmitting only structured information rather than full pixel data, it achieves over 99.9% bandwidth reduction compared to HEVC, as demonstrated on Olympic figure skating footage. The approach shifts computational load to the client and opens possibilities for volumetric avatar synthesis, multi‑view actor fusion, and personalized viewing experiences in a post‑codec era.
By Emanuele Artioli, Daniele Lorenzi, Shivi Vats, Farzad Tashtarian, Christian Timmerer
The paper introduces a GAN‑based semantic communication framework for image transmission in the Internet of Vehicles, aiming to overcome bandwidth and channel limitations. At the transmitter, a pyramid attention network extracts semantic label maps and a priority mechanism assigns weights to categories based on driving safety, guiding bit allocation and loss design. The receiver reconstructs images using a coarse‑to‑fine multi‑resolution generator, multi‑scale discriminator, temporal consistency, spatial pyramid pooling, and class‑aware convolutions, achieving high‑fidelity results with combined adversarial, feature‑matching, and perceptual losses. Experiments on Cityscapes demonstrate superior semantic segmentation accuracy and image quality compared to existing methods, with stable performance under AWGN and Rayleigh channels.
By Ruixing Ren, Shan Chen, Junhui Zhao, Xiaoke Sun
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.
arXiv:2505. 10946v3 Announce Type: replace-cross Abstract: Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation units across modalities.
By Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Robert Schober, Deniz G\"und\"uz
arXiv:2602. 12338v2 Announce Type: replace Abstract: Token Communications (TokenCom) has recently emerged as an effective new paradigm, where tokens are the unified units of multimodal communications and computations, enabling efficient digital semantic- and goal-oriented communications in future wireless networks.
By Farshad Zeinali, Mahdi Boloursaz Mashhadi, Rahim Tafazolli
arXiv:2606. 12858v1 Announce Type: cross Abstract: Conventional communication systems, including both separation-based coding and learning-based joint source-channel coding (JSCC), are typically designed under Shannon's rate-distortion theory.
By Tong Wu, Zhiyong Chen, Guo Lu, Li Song, Feng Yang, Meixia Tao, Wenjun Zhang
Ada-TokenCom is a rate‑adaptive token communication framework that uses large autoregressive models to compress and transmit tokens efficiently. It combines next‑token prediction with arithmetic coding, sending only the most informative tokens and letting the receiver generate the rest. A Lyapunov‑based algorithm dynamically adjusts compression and modulation to match changing network conditions, and simulations show it outperforms existing digital and deep joint source‑channel coding baselines.
By Zijun Zhang, Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Mehdi Bennis, Kaibin Huang
The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.
By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim