arXiv AI

JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications

arXiv:2606. 12858v1 Announce Type: cross Abstract: Conventional communication systems, including both separation-based coding and learning-based joint source-channel coding (JSCC), are typically designed under Shannon's rate-distortion theory.

arXiv Machine Learning
Sep 24

Advances in Diffusion-Based Generative Compression

This article reviews recent diffusion‑based methods for generative lossy image compression, highlighting how these techniques encode a source into an embedding and use a diffusion model to iteratively refine the reconstruction during decoding. It discusses the role of auxiliary entropy models for transmitting the embedding, explores the use of diffusion models for information transmission via channel simulation, and frames the discussion within rate‑distortion‑perception theory, common randomness, and inverse‑problem connections. The review also identifies open challenges in the field.

By Yibo Yang, Stephan Mandt
arXiv Machine Learning
Jun 30

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

arXiv:2604. 14603v2 Announce Type: replace-cross Abstract: The fundamental limit of natural signal compression has traditionally been characterized by classical rate-distortion (RD) theory through the tradeoff between coding rate and reconstruction distortion, while the rate-distortion-perception (RDP) framework introduces a divergence-based measure of perceptual quality as a modeling principle, leaving its theoretical origin unclear.

By Zijian Liang, Kai Niu, Changshuo Wang, Jin Xu, Ping Zhang
arXiv AI
6d ago

GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.

By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe
arXiv Computer Vision
Sep 22

Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement

The paper investigates why reconstruction quality in latent generative models does not always predict generative performance, attributing the issue to a mismatch between encoder-induced and generation-time latent distributions. It introduces Generation‑Aware Reconstruction (GAR), a method that perturbs encoder latents with noise and denoises them through the generative model before decoding, creating a continuous trajectory that reveals how the decoder behaves across latent spaces. The resulting GAR‑FID metric correlates strongly with generation FID, and using intermediate GAR latents for decoder adaptation consistently improves generative quality across different model scales.

By Xianghong Fang, Wenjie Shu, Tongda Xu, Wenlong Mou, Dehan Kong, Tim G. J. Rudner
Hugging Face Trending Papers
Aug 12

Generative Video Compression Based on Hierarchical Referencing

Diffusion-based generative video compression has emerged as a promising paradigm to improve perceptual quality, where latent frames are required to be encoded efficiently while serving as denoising conditions. However, existing methods neither carefully design reference and quality structures during latent coding nor account for the impact of frame-level quality variation on denoising procedure, which limits coding efficiency and aggravates artifact propagation during generative reconstruction.

arXiv Computer Vision
6d ago

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

arXiv:2609.31620v1 Announce Type: new Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...

By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv AI
Jul 8

Contrastive Predictive Coding with Compression for Enhanced Channel State Feedback in Wireless Networks

arXiv:2607. 05419v1 Announce Type: cross Abstract: Accurate and timely channel state information (CSI) is essential for next-generation wireless systems, yet existing works treat CSI compression and CSI prediction as separate problems, both in academia and in current 3GPP studies.

By Ahmed Y. Radwan, Hina Tabassum, Fahad Syed Muhammad, Matthew Baker