arXiv AI By Tong Wu, Zhiyong Chen, Guo Lu, Li Song, Feng Yang, Meixia Tao, Wenjun Zhang

JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications

Read the original on arXiv AI →

arXiv:2606. 12858v1 Announce Type: cross Abstract: Conventional communication systems, including both separation-based coding and learning-based joint source-channel coding (JSCC), are typically designed under Shannon's rate-distortion theory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 24

Advances in Diffusion-Based Generative Compression

This article reviews recent diffusion‑based methods for generative lossy image compression, highlighting how these techniques encode a source into an embedding and use a diffusion model to iteratively refine the reconstruction during decoding. It discusses the role of auxiliary entropy models for transmitting the embedding, explores the use of diffusion models for information transmission via channel simulation, and frames the discussion within rate‑distortion‑perception theory, common randomness, and inverse‑problem connections. The review also identifies open challenges in the field.

By Yibo Yang, Stephan Mandt
arXiv Machine Learning
Jun 30

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

arXiv:2604. 14603v2 Announce Type: replace-cross Abstract: The fundamental limit of natural signal compression has traditionally been characterized by classical rate-distortion (RD) theory through the tradeoff between coding rate and reconstruction distortion, while the rate-distortion-perception (RDP) framework introduces a divergence-based measure of perceptual quality as a modeling principle, leaving its theoretical origin unclear.

By Zijian Liang, Kai Niu, Changshuo Wang, Jin Xu, Ping Zhang
arXiv AI
6d ago

GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.

By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe
arXiv Computer Vision
Sep 22

Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement

The paper investigates why reconstruction quality in latent generative models does not always predict generative performance, attributing the issue to a mismatch between encoder-induced and generation-time latent distributions. It introduces Generation‑Aware Reconstruction (GAR), a method that perturbs encoder latents with noise and denoises them through the generative model before decoding, creating a continuous trajectory that reveals how the decoder behaves across latent spaces. The resulting GAR‑FID metric correlates strongly with generation FID, and using intermediate GAR latents for decoder adaptation consistently improves generative quality across different model scales.

By Xianghong Fang, Wenjie Shu, Tongda Xu, Wenlong Mou, Dehan Kong, Tim G. J. Rudner