arXiv:2607. 13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression.
By Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
S3VD is a new video deraining framework that leverages semantic guidance and spatio‑temporal scanning to improve performance over existing State Space Models such as Mamba. It introduces a Multi‑Scale Semantic Fusion module that uses DINOv2 priors to preserve 2D spatial semantics, and a Spatio‑Temporal Scanning Fusion module that incorporates a Decoupled‑Gating Mamba layer to better model intra‑ and inter‑frame correlations. Experiments on video deraining benchmarks show that S3VD achieves state‑of‑the‑art results, improving PSNR by an average of 0.84 dB over Mamba‑based baselines.
By Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu
arXiv:2608. 03822v1 Announce Type: cross Abstract: Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation.
By Zhang Weihui, Wang Ruizhi, Xu Hongye, Wang Huiqiong, Sun Li, Song Mingli
Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation. Although generative models provide a promising means of data augmentation, existing methods often yield implausible spatial layouts of flooded regions and distort scene structures.
arXiv:2607. 25321v1 Announce Type: new Abstract: Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity.
By Ruijie Su, Yuanzhi Liang, Xiaohua Xie, Jianhuang Lai
arXiv:2512.17323v2 Announce Type: replace
Abstract: Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video f...
By Jiyun Kong, Jun-Hyuk Kim, Jong-Seok Lee
arXiv:2503. 04500v3 Announce Type: replace-cross Abstract: Video understanding has largely relied on deep spatiotemporal architectures, including 3D convolutional networks and optical flow (OF) based models.
By Yu-Hsi Chen, Ching-Kai Lin, PingKong Huang, Chin-Tien Wu
arXiv:2609.36810v1 Announce Type: new
Abstract: Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determi...
By Renlong Wu, Guanqiao Wang, Xuan Shang, Yin Hanming, Xiaoxiao Sheng, Tianyu Huang, Hui Li, Wangmeng Zuo
arXiv:2608. 10544v1 Announce Type: cross Abstract: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations.
By Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
arXiv:2606. 06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding.
By Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen
arXiv:2608. 05981v1 Announce Type: new Abstract: High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains.
By Yichen Zhang, Yixiong Xiao, Congxi Xiao, Jingbo Zhou
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.