arXiv AI

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

arXiv:2607. 13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression.

arXiv Computer Vision
Aug 31

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

The paper introduces Parallel Tube Decoding (PTD), a generative approach for spatio‑temporal video grounding that splits the task into a temporal block and simultaneous time‑conditioned spatial blocks, eliminating token‑level and trajectory‑level dependencies. PTD uses Decoupled Block Attention to allow parallel spatial generation while maintaining shared video‑query context, and incorporates localization‑aware policy optimization for temporal boundaries and spatial geometry. Experiments on VidSTG show PTD cuts tube completion latency by 79× and boosts spatial decoding throughput by 92× compared to autoregressive decoding, while improving grounding accuracy and performing well on related tasks such as temporal grounding, VideoQA, and referring video object tracking.

By Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
arXiv Computer Vision
Sep 21

S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

S3VD is a new video deraining framework that leverages semantic guidance and spatio‑temporal scanning to improve performance over existing State Space Models such as Mamba. It introduces a Multi‑Scale Semantic Fusion module that uses DINOv2 priors to preserve 2D spatial semantics, and a Spatio‑Temporal Scanning Fusion module that incorporates a Decoupled‑Gating Mamba layer to better model intra‑ and inter‑frame correlations. Experiments on video deraining benchmarks show that S3VD achieves state‑of‑the‑art results, improving PSNR by an average of 0.84 dB over Mamba‑based baselines.

By Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu
arXiv Computer Vision
Sep 3

TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval

TAME is a CLIP‑based framework for Text‑Video Retrieval that incorporates temporal modeling through three key innovations: sparse Mixture‑of‑Experts layers with frame‑consistent routing, Frame‑Temporal tokens that aggregate cross‑frame information, and a Cross‑Temporal Interaction and Aggregation module for refining frame‑wise similarities. These components enable the model to capture both local visual patterns and long‑range temporal dependencies, leading to consistent performance gains over CLIP‑based baselines on multiple TVR benchmarks, including a 4.0 R@1 improvement on MSR‑VTT. The code is publicly available on GitHub.

By Uicheol Jung, Juyoung Hong, Hojung Kwon, Yukyung Choi
Hugging Face Trending Papers
Sep 2

TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval

TAME introduces a Temporal-Aware Mixture-of-Experts framework for Text-Video Retrieval that enhances CLIP-based models by incorporating frame-level structure and temporal relations. It adds sparse Mixture-of-Experts layers with frame-consistent routing, Frame-Temporal tokens for global cross-frame aggregation, and a Cross-Temporal Interaction and Aggregation module to refine sentence-video similarities. Experiments on multiple TVR benchmarks show consistent performance gains, such as a 4.0 R@1 improvement on MSR‑VTT over CLIP4Clip.

arXiv AI
6d ago

Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding

Pocket-STVG (P-STVG) is a lightweight cascade architecture for Spatio-Temporal Video Grounding that combines efficient pre‑trained components: a temporal‑aware video encoder based on MobileViCLIP, a spatial encoder‑decoder from MDETR, and a shared aligned text encoder. Temporal localization is achieved with a lightweight 1D U‑Net or a simple thresholding strategy, allowing the model to work in both weakly supervised and zero‑shot settings. With fewer than 90 M parameters, P-STVG matches or surpasses prior weakly supervised and zero‑shot methods while offering a more memory‑ and compute‑efficient pipeline for large‑scale video collections.

By Alberto Presta, Michal Byra, Grzegorz Stefa\'nski, Karol Szurkowski, Eryk Ko{\l}odziejczyk, Krzysztof Arendt
arXiv AI
Jun 12

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv Computer Vision
Aug 31

Training-Free Temporal Abstraction for General Video Understanding

The paper introduces STITCH, a training‑free method that partitions videos into semantically meaningful temporal chunks using a frozen video‑text backbone. By detecting changes in the embedding sequence of short video windows, STITCH produces reusable temporal abstractions that can be applied to multiple tasks such as event boundary detection, language‑based moment retrieval, and frame selection for vision‑language models. Experiments show that STITCH performs competitively with specialized methods while requiring no task‑specific training, especially when processing is limited to a few frames or tokens.

By Etienne Casanova, Sevan Brodjian, Pietro Perona