arXiv Computer Vision

ResARC: Residual-Aware AutoRegressive Coding for Ultra-Low Bitrate Image Compression

arXiv Computer Vision
1d ago

Rethinking Generative Image Compression at Extremely Low Bitrates

The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.

By Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu
arXiv Computer Vision
Sep 7

Multi-scale Image Representation Compression

The paper introduces MIRC, an overfitted image codec that quantizes and entropy‑codes all components—including latents, synthesis network, and entropy models—within a single end‑to‑end rate‑distortion framework inspired by NVRC. It adds a multi‑scale representation with cross‑stage parameter sharing to capture cross‑scale redundancy, yielding a 10.5 % BD‑rate saving over VVC on the CLIC2020 professional set. MIRC offers multiple configurations ranging from 1.2 to 2.9 kMAC per pixel, allowing decoding complexity to be tuned to deployment needs.

By Tianhao Peng, Ho Man Kwan, Fan Zhang, Shan Liu, David Bull
arXiv Computer Vision
Sep 4

Tree-Structured Vector Quantization For Efficient And Progressive Image Compression

Tree-VQ introduces a progressive tree‑structured vector quantization framework for learned image compression, organizing discrete codewords in a hierarchical binary tree where each latent token is represented by a routed root‑to‑leaf path. Every prefix of this path yields a valid quantized representation, enabling coarse reconstructions from shallow nodes and successive refinements from deeper nodes. The method incorporates a prefix‑compatible tree entropy model, rate‑aware refinement scheduling, and hierarchical prefix supervision to achieve efficient, low‑latency compression with superior perceptual quality and fewer parameters compared to existing approaches.

By Xinkun Wang, Tianyi Xu, Qingyu Luo, Mingming Ma, Changzhe Jiao, Fu Li, Yi Niu
Hugging Face Trending Papers
Sep 2

Multi-scale Image Representation Compression

The paper introduces MIRC, an overfitted image codec that quantizes all components—including latents, synthesis network, and entropy models—within a single rate‑distortion objective, following the neural video representation codec NVRC. It adds a multi‑scale representation with cross‑stage parameter sharing to capture cross‑scale redundancy, achieving a 10.5% BD‑rate saving over VVC on the CLIC2020 professional validation set. MIRC also offers configurable decoding complexity ranging from 1.2 to 2.9 kMAC per pixel, allowing deployment to match specific resource budgets.

arXiv AI
4d ago

GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.

By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe