The paper introduces Parallel Tube Decoding (PTD), a generative approach for spatio‑temporal video grounding that splits the task into a temporal block and simultaneous time‑conditioned spatial blocks, eliminating token‑level and trajectory‑level dependencies. PTD uses Decoupled Block Attention to allow parallel spatial generation while maintaining shared video‑query context, and incorporates localization‑aware policy optimization for temporal boundaries and spatial geometry. Experiments on VidSTG show PTD cuts tube completion latency by 79× and boosts spatial decoding throughput by 92× compared to autoregressive decoding, while improving grounding accuracy and performing well on related tasks such as temporal grounding, VideoQA, and referring video object tracking.
By Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
TempoGround is a vision‑language model–native framework for streaming visual grounding that detects cross‑frame object correspondence and explicitly models object presence states. It uses a curriculum prediction mechanism to resolve 2D instance association, predict object entry, continuation, or exit, decode 2D boxes, and lift them to 3D camera‑frame boxes. The approach is further refined with Streaming Grounding Reinforcement, which optimizes grounding, identity, and consistency rewards, and achieves significant improvements on multiple streaming visual grounding benchmarks.
By Leqian Ding, Junning Qiu, Manwen Yang, Yu Guo, Fei Wang
arXiv:2609.16722v1 Announce Type: new
Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates co...
By Haoyu Guo, Yuan Feng, Junlin Lv, Mingjun Xiao, S Kevin Zhou, Xike Xie
arXiv:2607. 13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression.
By Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
arXiv:2607. 02269v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG).
By Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma
WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining, using semantically coherent action events as the atomic unit of learning. Existing WAMs commonly initialize from multimodal or video foundation models and then optimize fixed-length action chunks conditioned directly on the current observation and instruction.