arXiv Computer Vision

Event-guided Neural Video Compression

The paper introduces Event-guided Neural Video Codec (ENVC), a neural video compression method that incorporates event streams—capturing brightness changes between frames—into both motion and frame coding stages. By using event-guided motion priors and event-conditioned predictors, ENVC improves RGB compression efficiency, achieving significant BD-rate savings across six benchmarks. The authors also synthesize paired RGB-event data for training and demonstrate that the gains persist on large-motion sequences, highlighting events as a valuable complementary modality for video coding.

arXiv Computer Vision
Sep 16

tcnerv:dual-domain temporal context modeling for implicit neural video compression

TCNeRV is a new implicit neural video compression method that models temporal context in both feature and embedding domains. Its multi‑scale temporal‑context fusion module injects gated historical features across decoder scales, while temporal embedding‑residual coding predicts and encodes only the residual of each content embedding. With about 3 million parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV‑Boost by 2.20 dB and reducing BD‑rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV respectively.

By Xuezhi Xiang, Yixin Zhao, Heqi Xiang, Jiayao Liu, Shanjun Zhang
arXiv Computer Vision
Aug 27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.

By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
arXiv Computer Vision
Sep 22

Dense-Prior-Guided Generative Video Compression

arXiv:2604.06655v2 Announce Type: replace Abstract: Diffusion-based generative video compression offers a promising paradigm for low-bitrate reconstruction, but existing keyframe-based controllable a...

By Ding Ding, Daowen Li, Yixin Gao, Ruixiao Dong, Kai Li, Ying Chen, Li Li
arXiv Computer Vision
2d ago

UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation

UniDynamics is a diffusion-based framework that generates future 4D dynamic scenes—comprising RGB, depth, and optical flow—from a single event-RGB pair, without needing long histories or control priors. It introduces an Event Latent Enhancement module to align event data into conditioning features and a Perceptual Dynamics Space within a multi-scale U‑Net to decouple and adaptively interact depth and flow, enforcing geometric and motion constraints for coherent predictions. Experiments on VKitti2 and DSEC show state‑of‑the‑art performance, producing high‑quality, temporally coherent, and 4D‑consistent predictions even under high‑speed motion blur.

By Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun
arXiv Machine Learning
Aug 7

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

arXiv:2608. 05728v1 Announce Type: cross Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval.

By Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan
Hugging Face Trending Papers
Jun 28

EvLIR: Learning Illumination Residuals from Ordered Events for Low-Light Image Enhancement

Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event cameras provide asynchronous brightness-change observations with high temporal resolution, but prior works often treat voxel channels as an unordered or static feature stack before fusion, rather than explicitly modeling their within-window temporal evolution, weakening the temporal evidence that makes events useful.