arXiv:2603.17546v2 Announce Type: replace
Abstract: Perceptual video compression leverages generative priors to reconstruct realistic textures and motions at low bitrates. However, existing perceptua...
By Daowen Li, Ruixiao Dong, Kai Li, Ying Chen, Ding Ding, Li Li
TCNeRV is a new implicit neural video compression method that models temporal context in both feature and embedding domains. Its multi‑scale temporal‑context fusion module injects gated historical features across decoder scales, while temporal embedding‑residual coding predicts and encodes only the residual of each content embedding. With about 3 million parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV‑Boost by 2.20 dB and reducing BD‑rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV respectively.
By Xuezhi Xiang, Yixin Zhao, Heqi Xiang, Jiayao Liu, Shanjun Zhang
arXiv:2602.19202v3 Announce Type: replace
Abstract: Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensit...
By Gang Xu, Zhiyu Zhu, Junhui Hou
arXiv:2608.28429v1 Announce Type: new
Abstract: Event cameras generate asynchronous, sparse data streams with microsecond temporal resolution, but in moderate-to-high motion scenes they can produce a...
By Zahra Rezaee, Catarina Brites, Jo\~ao Ascenso
arXiv:2606. 02569v1 Announce Type: cross Abstract: Video is temporally redundant: adjacent frames usually share most objects, background, and layout.
By Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si, Chenglin Li, Shuai Dong, Kele Shao, Ruilin Li, Dianyi Wang, Nan Duan, Jiaqi Wang
LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.
By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
arXiv:2604.06655v2 Announce Type: replace
Abstract: Diffusion-based generative video compression offers a promising paradigm for low-bitrate reconstruction, but existing keyframe-based controllable a...
By Ding Ding, Daowen Li, Yixin Gao, Ruixiao Dong, Kai Li, Ying Chen, Li Li
UniDynamics is a diffusion-based framework that generates future 4D dynamic scenes—comprising RGB, depth, and optical flow—from a single event-RGB pair, without needing long histories or control priors. It introduces an Event Latent Enhancement module to align event data into conditioning features and a Perceptual Dynamics Space within a multi-scale U‑Net to decouple and adaptively interact depth and flow, enforcing geometric and motion constraints for coherent predictions. Experiments on VKitti2 and DSEC show state‑of‑the‑art performance, producing high‑quality, temporally coherent, and 4D‑consistent predictions even under high‑speed motion blur.
By Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun
arXiv:2608. 05728v1 Announce Type: cross Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval.
By Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan
arXiv:2605. 00271v3 Announce Type: replace-cross Abstract: Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and robustness to extreme lighting.
By Vincenzo Polizzi, David B. Lindell, Jonathan Kelly
Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event cameras provide asynchronous brightness-change observations with high temporal resolution, but prior works often treat voxel channels as an unordered or static feature stack before fusion, rather than explicitly modeling their within-window temporal evolution, weakening the temporal evidence that makes events useful.
Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Training (TTT) resampler...