The paper introduces ACV-Gate, an adaptive framework for generative image communication that selectively evaluates candidate semantic tokens to reduce encoder computation while maintaining reconstruction quality. ACV-Gate uses a set-aware student model trained on terminal advantages and regrets to rank tokens, and a selective refinement mechanism that limits exact evaluations to a bounded candidate set, controlled by cost-based thresholds. Experiments on CIFAR-10, STL-10, and 384×384 images show that ACV-Gate improves PSNR by up to 0.636 dB at 0.20 bpp and reduces candidate evaluations to about 27.6% of the exact-full method.
By Qinglei Qi, Zhihe Liang, Fengzhan Jing, Shenao Zhu, Lei Zhang, Chenyang Zhang, Shuqing He, Jia Guo
The paper introduces PRESLEY, an end‑to‑end video streaming pipeline that uses generative AI layers to selectively degrade and restore less important regions of a frame. By replacing destructive block removal with adaptive in‑place degradation and signaling block strength via a side channel, PRESLEY achieves significant bitrate savings and improved background quality compared to its predecessor and pristine baselines. The authors also analyze the theoretical headroom of this architecture, quantifying remaining cost‑axis headroom and modeling post‑restoration damage to guide future rate‑distortion‑restoration selection.
By Emanuele Artioli, Farzad Tashtarian, Christian Timmerer
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid.
The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.
By Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu
The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.
By Cheng Yuan, Jiawei Shao, Xuelong Li
arXiv:2608. 08698v1 Announce Type: new Abstract: Video token communication represents video content as discrete tokens that differ in their importance to reconstruction and exhibit temporal dependencies.
By Bingyan Xie, Yongjeong Oh, Zihan Chen, Jihong Park, Yongpeng Wu, Wenjun Zhang