arXiv Computer Vision

Projection-Aware End-to-End Learned Video Compression for 360-Degree Video

The study examines how different two‑dimensional projections of spherical 360‑degree video affect end‑to‑end neural compression. Seven JVET‑360Lib formats were evaluated using a scale‑space flow model on standard test sequences, with performance measured by PSNR, spherical PSNR, weighted spherical PSNR, and BD‑rate. Results show that equirectangular and padded equirectangular projections yield the best compression efficiency with the neural model, while cubemap‑based formats excel with conventional HM‑16.16 codecs, highlighting that projection choice is codec‑dependent.

arXiv Computer Vision
Sep 7

Scalable Neural Video Representation Compression

Scalable Neural Video Representation Compression (S-NVRC) introduces a scalable implicit neural representation (INR) video codec that supports fine-grained bitrate and decoding‑complexity scalability from a single embedded bitstream. It uses a coarse‑to‑fine prefix for feature grids and a nested prefix for network layers, enabling a wide range of operating points while maintaining a single encoding. On the UVG dataset, S‑NVRC outperforms SHM 12.4 and multi‑layer VTM‑20.0 by 43.7 % and 5.6 % in BD‑rate, respectively, and offers flexible complexity scalability.

By Tianhao Peng, Ho Man Kwan, Fan Zhang, Shan Liu, David Bull
arXiv Computer Vision
2d ago

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.

By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.

arXiv AI
Sep 4

LRConv-NeRV: Low Rank Convolution for Efficient Neural Video Compression

LRConv-NeRV introduces low‑rank separable convolutions into the NeRV neural video decoder, replacing selected dense 3x3 layers to reduce computational load and memory usage. By applying low‑rank factorization progressively from the largest to earlier decoder stages, the method offers controllable trade‑offs between reconstruction quality and efficiency. Experiments show that applying LRConv only to the final decoder stage cuts decoder complexity by 68% and model size by 9.3% with negligible quality loss, while INT8 quantization preserves performance close to the dense baseline.

By Tamer Shanableh
arXiv Computer Vision
Aug 28

KISS-GS: 3D Gaussian Splatting Compression Kept Simple

KISS-GS is a modular compression pipeline for 3D Gaussian Splatting (3DGS) scenes that separates compression from training. It first compacts a vanilla 3DGS scene by 15.7× using state‑of‑the‑art pruning, then encodes the result into the SOG‑XT image‑based format, achieving an additional 6.6× reduction. Optional encoding‑aware fine‑tuning can further cut the size by 2.2×, yielding total reductions of 85× to 319× on standard benchmarks while enabling web‑native decoding.

By Wieland Morgenstern, Friedrich Elias Branschke, Florian Fleischmann, Adrian Szatmari, Paul Schlack, Florian Barthel, Peter Eisert, Anna Hilsmann
arXiv AI
Sep 2

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

The paper investigates whether large language models (LLMs) can design video coding tools, focusing on the Planar mode used in video coding standards. Using a generation-and-evaluation loop, the LLM generates new Planar predictors, which are then tested in the Fraunhofer Versatile Video Encoder (VVenC) and the Enhanced Compression Model (ECM). Results show that the LLM-generated mode can outperform the conventional Planar mode, achieving a 0.18% bitrate saving with a 0.4% complexity increase, and that similar gains are possible when integrating the new predictor into ECM under low‑resolution settings.

By Yingwen Zhang, Meng Wang, Liqiang He, Shiqi Wang
arXiv AI
Sep 17

GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media

GenStream is a semantic streaming framework that replaces dense video frames with compact metadata—skeletal keypoints, camera parameters, and a static 3D background model—to enable generative reconstruction of human figures on the client side. By transmitting only structured information rather than full pixel data, it achieves over 99.9% bandwidth reduction compared to HEVC, as demonstrated on Olympic figure skating footage. The approach shifts computational load to the client and opens possibilities for volumetric avatar synthesis, multi‑view actor fusion, and personalized viewing experiences in a post‑codec era.

By Emanuele Artioli, Daniele Lorenzi, Shivi Vats, Farzad Tashtarian, Christian Timmerer