arXiv AI

FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring

arXiv:2512. 04390v2 Announce Type: replace-cross Abstract: Joint video super-resolution and deblurring (VSRDB) requires both efficient long-range temporal modeling and robustness to frame-wise exposure-duration variation, which changes the extent of motion blur across video frames.

arXiv AI
1d ago

PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video

PickMoment is a continuous‑time model that learns to predict the interval‑mean blur over arbitrary sub‑intervals of a camera exposure, unifying single‑image deblurring, blur‑to‑video generation, and continuous‑time pick‑a‑moment recovery. It is trained with three supervisions derived from the blur integral: an empirical reconstruction loss, an additivity loss for self‑consistency, and a sharp‑frame loss at zero interval. The model achieves state‑of‑the‑art performance on GoPro and HIDE for generative deblurring, competitive results on RealBlur, and the highest per‑frame fidelity on GoPro‑7 blur‑to‑video, all in a single forward pass.

By Junseong Shin, Hyeonsu Jo, Daehyun Kim, Tae Hyun Kim
arXiv Computer Vision
Sep 10

SloMoDeblur: A Large-Scale Smartphone Image Deblurring Dataset

arXiv:2506.19445v5 Announce Type: replace Abstract: Motion blur remains one of the most common and visually disruptive degradations in real-world smartphone imaging, yet existing deblurring benchmark...

By Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder, Abdul Mohaimen Al Radi, Sudipto Das Sukanto, Afia Lubaina, Md. Mosaddek Khan
arXiv AI
Jul 24

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring

arXiv:2607. 20628v1 Announce Type: cross Abstract: Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction.

By Renbiao Jin, Mingxin Yang, Yutian Chen, Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue
arXiv AI
Aug 28

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.

By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
arXiv Computer Vision
Aug 31

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.

By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
arXiv Computer Vision
Aug 27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.

By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
arXiv AI
5d ago

GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.

By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe
arXiv AI
Sep 4

LRConv-NeRV: Low Rank Convolution for Efficient Neural Video Compression

LRConv-NeRV introduces low‑rank separable convolutions into the NeRV neural video decoder, replacing selected dense 3x3 layers to reduce computational load and memory usage. By applying low‑rank factorization progressively from the largest to earlier decoder stages, the method offers controllable trade‑offs between reconstruction quality and efficiency. Experiments show that applying LRConv only to the final decoder stage cuts decoder complexity by 68% and model size by 9.3% with negligible quality loss, while INT8 quantization preserves performance close to the dense baseline.

By Tamer Shanableh
arXiv Machine Learning
Aug 7

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

arXiv:2608. 05728v1 Announce Type: cross Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval.

By Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan