arXiv Computer Vision By Jiyun Kong, Jungwoo Kim, Enes Eray Demirtas, Touradj Ebrahimi, Jong-Seok Lee

Event-guided Neural Video Compression

Read the original on arXiv Computer Vision →

The paper introduces Event-guided Neural Video Codec (ENVC), a neural video compression method that incorporates event streams—capturing brightness changes between frames—into both motion and frame coding stages. By using event-guided motion priors and event-conditioned predictors, ENVC improves RGB compression efficiency, achieving significant BD-rate savings across six benchmarks. The authors also synthesize paired RGB-event data for training and demonstrate that the gains persist on large-motion sequences, highlighting events as a valuable complementary modality for video coding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 16

tcnerv:dual-domain temporal context modeling for implicit neural video compression

TCNeRV is a new implicit neural video compression method that models temporal context in both feature and embedding domains. Its multi‑scale temporal‑context fusion module injects gated historical features across decoder scales, while temporal embedding‑residual coding predicts and encodes only the residual of each content embedding. With about 3 million parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV‑Boost by 2.20 dB and reducing BD‑rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV respectively.

By Xuezhi Xiang, Yixin Zhao, Heqi Xiang, Jiayao Liu, Shanjun Zhang
arXiv Computer Vision
Aug 27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.

By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny