arXiv Computer Vision By Niloofar Maani

Projection-Aware End-to-End Learned Video Compression for 360-Degree Video

Read the original on arXiv Computer Vision →

The study examines how different two‑dimensional projections of spherical 360‑degree video affect end‑to‑end neural compression. Seven JVET‑360Lib formats were evaluated using a scale‑space flow model on standard test sequences, with performance measured by PSNR, spherical PSNR, weighted spherical PSNR, and BD‑rate. Results show that equirectangular and padded equirectangular projections yield the best compression efficiency with the neural model, while cubemap‑based formats excel with conventional HM‑16.16 codecs, highlighting that projection choice is codec‑dependent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 7

Scalable Neural Video Representation Compression

Scalable Neural Video Representation Compression (S-NVRC) introduces a scalable implicit neural representation (INR) video codec that supports fine-grained bitrate and decoding‑complexity scalability from a single embedded bitstream. It uses a coarse‑to‑fine prefix for feature grids and a nested prefix for network layers, enabling a wide range of operating points while maintaining a single encoding. On the UVG dataset, S‑NVRC outperforms SHM 12.4 and multi‑layer VTM‑20.0 by 43.7 % and 5.6 % in BD‑rate, respectively, and offers flexible complexity scalability.

By Tianhao Peng, Ho Man Kwan, Fan Zhang, Shan Liu, David Bull
arXiv Computer Vision
2d ago

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.

By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.