arXiv AI By Alberto Presta, Michal Byra, Grzegorz Stefa\'nski, Karol Szurkowski, Eryk Ko{\l}odziejczyk, Krzysztof Arendt

Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding

Read the original on arXiv AI →

Pocket-STVG (P-STVG) is a lightweight cascade architecture for Spatio-Temporal Video Grounding that combines efficient pre‑trained components: a temporal‑aware video encoder based on MobileViCLIP, a spatial encoder‑decoder from MDETR, and a shared aligned text encoder. Temporal localization is achieved with a lightweight 1D U‑Net or a simple thresholding strategy, allowing the model to work in both weakly supervised and zero‑shot settings. With fewer than 90 M parameters, P-STVG matches or surpasses prior weakly supervised and zero‑shot methods while offering a more memory‑ and compute‑efficient pipeline for large‑scale video collections.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 31

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

The paper introduces Parallel Tube Decoding (PTD), a generative approach for spatio‑temporal video grounding that splits the task into a temporal block and simultaneous time‑conditioned spatial blocks, eliminating token‑level and trajectory‑level dependencies. PTD uses Decoupled Block Attention to allow parallel spatial generation while maintaining shared video‑query context, and incorporates localization‑aware policy optimization for temporal boundaries and spatial geometry. Experiments on VidSTG show PTD cuts tube completion latency by 79× and boosts spatial decoding throughput by 92× compared to autoregressive decoding, while improving grounding accuracy and performing well on related tasks such as temporal grounding, VideoQA, and referring video object tracking.

By Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
arXiv AI
Jul 29

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.

By Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin