arXiv AI

WS-NeRF: A Mamba-Driven World-State-Aware Adaptive Deblurring Neural Radiance Field

arXiv AI
Jul 24

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring

arXiv:2607. 20628v1 Announce Type: cross Abstract: Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction.

By Renbiao Jin, Mingxin Yang, Yutian Chen, Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue
arXiv AI
2d ago

PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video

PickMoment is a continuous‑time model that learns to predict the interval‑mean blur over arbitrary sub‑intervals of a camera exposure, unifying single‑image deblurring, blur‑to‑video generation, and continuous‑time pick‑a‑moment recovery. It is trained with three supervisions derived from the blur integral: an empirical reconstruction loss, an additivity loss for self‑consistency, and a sharp‑frame loss at zero interval. The model achieves state‑of‑the‑art performance on GoPro and HIDE for generative deblurring, competitive results on RealBlur, and the highest per‑frame fidelity on GoPro‑7 blur‑to‑video, all in a single forward pass.

By Junseong Shin, Hyeonsu Jo, Daehyun Kim, Tae Hyun Kim
arXiv Computer Vision
Sep 10

SloMoDeblur: A Large-Scale Smartphone Image Deblurring Dataset

arXiv:2506.19445v5 Announce Type: replace Abstract: Motion blur remains one of the most common and visually disruptive degradations in real-world smartphone imaging, yet existing deblurring benchmark...

By Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder, Abdul Mohaimen Al Radi, Sudipto Das Sukanto, Afia Lubaina, Md. Mosaddek Khan
arXiv Computer Vision
Sep 23

NaCR: Visual Localization via NeRF-aided Camera Ray Regression

NaCR: Visual Localization via NeRF-aided Camera Ray Regression proposes a unified framework that integrates Neural Radiance Fields (NeRF) with Camera Ray Regression (CRR) to improve visual localization accuracy. The method enhances the CRR baseline with three simple improvements, augments training data by synthesizing novel views from a pre‑trained NeRF, and employs a closed‑loop supervision pipeline that back‑propagates photometric rendering errors to refine predicted camera rays. A two‑stage training curriculum ensures stable convergence, and experiments on indoor and outdoor benchmarks show competitive accuracy with validated component efficacy.

By Yesheng Zhang, Xiang Dai, Xu Zhao, Chongyang Zhang
arXiv Computer Vision
Aug 25

Controllable blind deblurring with diffusion models

The paper introduces SuperSharpen, a diffusion-based method for blind deblurring in professional photography that can invert unknown isotropic blur without knowing the degradation kernel. It offers explicit control over restoration strength via a blur measure and compares two conditioning strategies: a ControlNet-style adapter on a frozen backbone and full finetuning of the diffusion prior. Experiments on synthetic and real-world blur show that finetuning yields higher fidelity with fewer hallucinated details, improving perceptual quality and controllable restoration strength.

By Imane Si Salah, Emile Cribelier, Thomas Veit, Wolf Hauser, Arthur Leclaire
arXiv Computer Vision
Aug 28

Loop-Mamba: A Loop Mamba with Degradation-Aware and Shared Memory for Old Photo Restoration

Loop‑Mamba is a lightweight, loop‑based state‑space framework designed for restoring old photographs that suffer from multiple degradations such as scratches, cracks, fading, blur, noise, and missing regions. It models restoration as progressive state evolution, using a Semantic‑Guided Degradation Estimator to predict local degradation maps and global scores, and a Shared Structural Memory Mamba to maintain a persistent restoration state across iterations. The method employs first‑order state recursion and a multi‑directional scanning strategy to reduce gradient dilution and computational overhead, and introduces the Old Photo Damage Recovery Score (ODRS) to evaluate both degradation recovery and structural reconstruction, achieving superior performance on the SynOld benchmark.

By Runci Bai, Yucheng Xin, Pu Wang, Yongcong Wang, Chen Wu, Dianjie Lu, Guijuan Zhang, Pengwen Dai, Guangwei Gao, Siyuan Yao, Zhuoran Zheng
arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang
Hugging Face Trending Papers
Aug 11

CasDeblurGS: Cascaded 2D-to-3D Multi-View Consistency for 3D Gaussian Splatting from Two Blurry Images

Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and motion blur. Although neural rendering has advanced sparse-view synthesis, existing blur-aware methods typically require substantial multi-view redundancy, accurate camera poses, or costly per-scene optimization.