arXiv:2607. 13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression.
By Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
S3VD is a new video deraining framework that leverages semantic guidance and spatio‑temporal scanning to improve performance over existing State Space Models such as Mamba. It introduces a Multi‑Scale Semantic Fusion module that uses DINOv2 priors to preserve 2D spatial semantics, and a Spatio‑Temporal Scanning Fusion module that incorporates a Decoupled‑Gating Mamba layer to better model intra‑ and inter‑frame correlations. Experiments on video deraining benchmarks show that S3VD achieves state‑of‑the‑art results, improving PSNR by an average of 0.84 dB over Mamba‑based baselines.
By Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu
arXiv:2608. 03822v1 Announce Type: cross Abstract: Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation.
By Zhang Weihui, Wang Ruizhi, Xu Hongye, Wang Huiqiong, Sun Li, Song Mingli
Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation. Although generative models provide a promising means of data augmentation, existing methods often yield implausible spatial layouts of flooded regions and distort scene structures.
arXiv:2607. 25321v1 Announce Type: new Abstract: Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity.
By Ruijie Su, Yuanzhi Liang, Xiaohua Xie, Jianhuang Lai
arXiv:2512.17323v2 Announce Type: replace
Abstract: Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video f...
By Jiyun Kong, Jun-Hyuk Kim, Jong-Seok Lee