arXiv AI

SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance

arXiv:2602. 21819v3 Announce Type: replace-cross Abstract: Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception.

arXiv Computer Vision
Sep 21

S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

S3VD is a new video deraining framework that leverages semantic guidance and spatio‑temporal scanning to improve performance over existing State Space Models such as Mamba. It introduces a Multi‑Scale Semantic Fusion module that uses DINOv2 priors to preserve 2D spatial semantics, and a Spatio‑Temporal Scanning Fusion module that incorporates a Decoupled‑Gating Mamba layer to better model intra‑ and inter‑frame correlations. Experiments on video deraining benchmarks show that S3VD achieves state‑of‑the‑art results, improving PSNR by an average of 0.84 dB over Mamba‑based baselines.

By Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu
arXiv AI
Sep 7

ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding

ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding introduces a model‑agnostic framework that adaptively aligns EEG signals with visual semantics. It replaces fixed visual or textual anchors with EEG‑aware class‑level contrastive supervision and employs structure‑consistent interpolation to preserve channel‑wise and temporal importance. Across multiple evaluation settings—including subject‑dependent, subject‑independent, strict cross‑subject transfer, and continual adaptation—ProCA delivers significant performance gains, achieving relative Top‑1 improvements ranging from 7.4% to 28.1%.

By Kanglei Zhou, Chunyan Lan, Dongyang Li, Jun Zhu, Liyuan Wang
Hugging Face Trending Papers
Sep 2

TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval

TAME introduces a Temporal-Aware Mixture-of-Experts framework for Text-Video Retrieval that enhances CLIP-based models by incorporating frame-level structure and temporal relations. It adds sparse Mixture-of-Experts layers with frame-consistent routing, Frame-Temporal tokens for global cross-frame aggregation, and a Cross-Temporal Interaction and Aggregation module to refine sentence-video similarities. Experiments on multiple TVR benchmarks show consistent performance gains, such as a 4.0 R@1 improvement on MSR‑VTT over CLIP4Clip.