arXiv AI

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

arXiv:2608. 16192v1 Announce Type: new Abstract: Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver.

arXiv AI
6d ago

Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication

The paper introduces ACV-Gate, an adaptive framework for generative image communication that selectively evaluates candidate semantic tokens to reduce encoder computation while maintaining reconstruction quality. ACV-Gate uses a set-aware student model trained on terminal advantages and regrets to rank tokens, and a selective refinement mechanism that limits exact evaluations to a bounded candidate set, controlled by cost-based thresholds. Experiments on CIFAR-10, STL-10, and 384×384 images show that ACV-Gate improves PSNR by up to 0.636 dB at 0.20 bpp and reduces candidate evaluations to about 27.6% of the exact-full method.

By Qinglei Qi, Zhihe Liang, Fengzhan Jing, Shenao Zhu, Lei Zhang, Chenyang Zhang, Shuqing He, Jia Guo
arXiv AI
Sep 18

Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers

The paper introduces PRESLEY, an end‑to‑end video streaming pipeline that uses generative AI layers to selectively degrade and restore less important regions of a frame. By replacing destructive block removal with adaptive in‑place degradation and signaling block strength via a side channel, PRESLEY achieves significant bitrate savings and improved background quality compared to its predecessor and pristine baselines. The authors also analyze the theoretical headroom of this architecture, quantifying remaining cost‑axis headroom and modeling post‑restoration damage to guide future rate‑distortion‑restoration selection.

By Emanuele Artioli, Farzad Tashtarian, Christian Timmerer
arXiv Computer Vision
3d ago

Rethinking Generative Image Compression at Extremely Low Bitrates

The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.

By Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu
arXiv Computer Vision
Sep 24

Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality

The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.

By Cheng Yuan, Jiawei Shao, Xuelong Li
arXiv AI
6d ago

GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.

By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe
Hugging Face Trending Papers
Jun 17

GB-LSR: A Fast Local Spectral Image Representation with a Single Global Bandwidth for Continuous Reconstruction and Super-Resolution

We present GB-LSR (Global-Bandwidth Local Spectral Representation), a fixed-grid local spectral representation for continuous image reconstruction. The image domain is partitioned into non-overlapping square patches, each carrying coefficients for a truncated Fourier basis predicted from shared convolutional-encoder features.

arXiv Machine Learning
Sep 10

Token Encoding for Semantic Recovery

The paper introduces TokCode, a token encoding framework that enhances robustness in generative semantic communication by restructuring redundancy in the semantic domain. TokCode leverages a lightweight adapter to transform a large language model into a token encoder, avoiding the need for a dedicated deep model. A channel-quality-aware distillation method (CADET) trains the adapter across diverse erasure rates, producing a reconfigurable low‑rank adapter that enables efficient reinforcement learning and achieves significant improvements in image similarity over existing receiver‑side recovery benchmarks.

By Jingzhi Hu, Ouya Wang, Geoffrey Ye Li