arXiv Computer Vision

FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining

FluidRain is a lightweight video deraining model that leverages a divergence‑free rain flow field to guide Loop‑in‑Loop attention across scales and neighboring frames, eliminating the need for explicit motion alignment. By projecting estimated rain‑flow onto a divergence‑free subspace, the method steers window attention along rain streaks, enabling efficient temporal aggregation with only 0.80 M parameters. Experiments on four benchmarks demonstrate competitive performance against larger models, and the authors introduce a new RainSyn‑Gust dataset and a physics‑based no‑reference metric for evaluating real‑rain removal.

arXiv Computer Vision
Sep 21

S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

S3VD is a new video deraining framework that leverages semantic guidance and spatio‑temporal scanning to improve performance over existing State Space Models such as Mamba. It introduces a Multi‑Scale Semantic Fusion module that uses DINOv2 priors to preserve 2D spatial semantics, and a Spatio‑Temporal Scanning Fusion module that incorporates a Decoupled‑Gating Mamba layer to better model intra‑ and inter‑frame correlations. Experiments on video deraining benchmarks show that S3VD achieves state‑of‑the‑art results, improving PSNR by an average of 0.84 dB over Mamba‑based baselines.

By Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu
arXiv AI
Jul 29

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

arXiv:2607. 25321v1 Announce Type: new Abstract: Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity.

By Ruijie Su, Yuanzhi Liang, Xiaohua Xie, Jianhuang Lai
arXiv Computer Vision
4d ago

MeteoVerse: Unified Weather-Controllable Video World Model

arXiv:2609.36810v1 Announce Type: new Abstract: Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determi...

By Renlong Wu, Guanqiao Wang, Xuan Shang, Yin Hanming, Xiaoxiao Sheng, Tianyu Huang, Hui Li, Wangmeng Zuo
arXiv AI
Aug 12

Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration

arXiv:2608. 10544v1 Announce Type: cross Abstract: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations.

By Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
Hugging Face Trending Papers
Jul 14

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.