arXiv Machine Learning

Semantic-Aware Joint Source-Channel Optimization for Encoder-Agnostic Digital Video Communication

arXiv AI
Sep 4

LRConv-NeRV: Low Rank Convolution for Efficient Neural Video Compression

LRConv-NeRV introduces low‑rank separable convolutions into the NeRV neural video decoder, replacing selected dense 3x3 layers to reduce computational load and memory usage. By applying low‑rank factorization progressively from the largest to earlier decoder stages, the method offers controllable trade‑offs between reconstruction quality and efficiency. Experiments show that applying LRConv only to the final decoder stage cuts decoder complexity by 68% and model size by 9.3% with negligible quality loss, while INT8 quantization preserves performance close to the dense baseline.

By Tamer Shanableh
arXiv Computer Vision
Sep 24

Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality

The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.

By Cheng Yuan, Jiawei Shao, Xuelong Li
arXiv Computer Vision
Sep 16

High-Fidelity Video Quality Assessment with VQA-Specific Saliency

High-Fidelity Video Quality Assessment (HFVQA) is a new framework that uses fixed-size spatio‑temporal patches across multiple scales, including the original resolution, to preserve low‑level quality cues and semantic context. It incorporates a lightweight auxiliary network that learns VQA‑specific saliency directly from quality supervision, enabling the model to focus on the most important spatio‑temporal regions. By combining high‑fidelity cues with task‑specific saliency, HFVQA achieves state‑of‑the‑art performance on standard no‑reference VQA benchmarks while processing only about 12% of the candidate patches, making it computationally efficient.

By Hakan Emre Gedik, Shashank Gupta, Alan Bovik
Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.

arXiv Computer Vision
Sep 7

Scalable Neural Video Representation Compression

Scalable Neural Video Representation Compression (S-NVRC) introduces a scalable implicit neural representation (INR) video codec that supports fine-grained bitrate and decoding‑complexity scalability from a single embedded bitstream. It uses a coarse‑to‑fine prefix for feature grids and a nested prefix for network layers, enabling a wide range of operating points while maintaining a single encoding. On the UVG dataset, S‑NVRC outperforms SHM 12.4 and multi‑layer VTM‑20.0 by 43.7 % and 5.6 % in BD‑rate, respectively, and offers flexible complexity scalability.

By Tianhao Peng, Ho Man Kwan, Fan Zhang, Shan Liu, David Bull
arXiv AI
Jul 29

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

arXiv:2607. 25669v1 Announce Type: new Abstract: Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs.

By Haoyang Huang, Wenjie Huang, Tianqi Xu, Hongyaoxing Gu, Kang Tan, Yikai Fu, Yuhao Shen, Tianyu Liu, Baolin Zhang, Jun Zhang, Xinyi Hu, Jun Dai, Shuang Ge, Lei Chen, Yue Li, Mingchen Wang, Meng Zhang
arXiv Computer Vision
4d ago

RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

RelayVSR introduces a streaming video super‑resolution framework that combines a large generative model, which produces reference latents for sparse keyframes, with a lightweight Dual‑Memory Video Transformer that super‑resolves every frame using these references and low‑resolution input. The method employs Video‑Aware Reference Optimization (VARO), a reinforcement‑learning strategy that optimizes both system‑level video quality and reference‑level keyframe fidelity, outperforming direct joint training. On 1080p video, RelayVSR achieves 29.29 FPS with modest GPU memory usage, significantly faster and more efficient than the FlashVSR‑Tiny baseline.

By Xijun Wang, Xin Li, Zirui Lang, Suhang Yao, Haoran Li, Zhibo Chen
arXiv Machine Learning
Sep 16

Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission

The paper introduces a semantic‑aware multi‑level neural video codec designed for low‑latency, task‑oriented video transmission over unreliable channels. It builds on the real‑time DCVC‑RT codec by partitioning encoded representations into packets of varying semantic and feature importance, assigning them to priority streams, and employing an error‑resilient entropy model that removes inter‑packet dependencies. Experiments demonstrate that this framework improves robustness against packet erasures, achieving graceful degradation in less important regions while preserving task‑relevant visual content.

By Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du, Harish Viswanathan
arXiv AI
6d ago

GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.

By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe