arXiv Computer Vision

S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

S3VD is a new video deraining framework that leverages semantic guidance and spatio‑temporal scanning to improve performance over existing State Space Models such as Mamba. It introduces a Multi‑Scale Semantic Fusion module that uses DINOv2 priors to preserve 2D spatial semantics, and a Spatio‑Temporal Scanning Fusion module that incorporates a Decoupled‑Gating Mamba layer to better model intra‑ and inter‑frame correlations. Experiments on video deraining benchmarks show that S3VD achieves state‑of‑the‑art results, improving PSNR by an average of 0.84 dB over Mamba‑based baselines.

arXiv Computer Vision
Sep 25

FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining

FluidRain is a lightweight video deraining model that leverages a divergence‑free rain flow field to guide Loop‑in‑Loop attention across scales and neighboring frames, eliminating the need for explicit motion alignment. By projecting estimated rain‑flow onto a divergence‑free subspace, the method steers window attention along rain streaks, enabling efficient temporal aggregation with only 0.80 M parameters. Experiments on four benchmarks demonstrate competitive performance against larger models, and the authors introduce a new RainSyn‑Gust dataset and a physics‑based no‑reference metric for evaluating real‑rain removal.

By Pu Wang, Yongcong Wang, Wenhao Li, Xiang Chen, Guangwei Gao, Jinshan Pan, Siyuan Yao, Shujun Fu, Zhuoran Zheng
arXiv Computer Vision
Aug 24

Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving

The paper introduces Driving with DINO (DwD), a framework that uses Vision Foundation Module (VFM) features to bridge simulation and real-world domains for autonomous driving video generation. It addresses the consistency‑realism dilemma by projecting VFM features onto a principal subspace, dropping high‑frequency texture elements, and applying a Random Channel Tail Drop to preserve structural detail. Additional components— a learnable Spatial Alignment Module and a Causal Temporal Aggregator— enhance control precision, spatial alignment, and temporal stability, reducing motion blur and ensuring realistic, consistent outputs.

By Xuyang Chen, Conglang Zhang, Chuanheng Fu, Zihao Yang, Kaixuan Zhou, Yizhi Zhang, Yanfeng Zhang, Mingwei Sun, Zhen Dong, Xiaoxiao Long, Zengmao Wang, Liqiu Meng