arXiv AI

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

arXiv:2607. 02551v1 Announce Type: cross Abstract: Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception.

arXiv AI
Sep 10

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

TimeBlind is a diagnostic benchmark designed to evaluate fine‑grained spatio‑temporal compositionality in video large language models (LLMs). It categorizes temporal understanding into three levels—atomic event recognition, event property characterization, and reasoning about event interdependencies—and uses a minimal‑pairs paradigm where video pairs share identical static content but differ only in temporal structure. Across 20 state‑of‑the‑art MLLMs tested on 600 curated instances, the best model achieved only 48.2% instance accuracy, far below human performance of 98.2%, highlighting a reliance on static visual shortcuts rather than true temporal reasoning.

By Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius
arXiv AI
Jun 2

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

arXiv:2606. 02522v1 Announce Type: cross Abstract: Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored.

By Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang, Yan Li, Xin Li, Haoyu Cao, Xing Sun, Shaofeng Zhang, Xu Yang, Zhihang Zhong, Xue Yang
arXiv Computer Vision
3d ago

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.

By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
arXiv Computation and Language
Sep 25

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND is a new benchmark that tests multimodal large language models’ ability to track objects, their states, and relationships over time in videos. It evaluates intermediate reasoning by breaking queries into sub‑questions and uses Faithful Accuracy to ensure all parts of an answer are correct. The authors also propose an object‑centric framework that builds structured trajectories and shows reduced hallucinations and better temporal consistency compared to existing models.

By Thong Nguyen, Tri Cao, Khoi Le, Cong-Duy Nguyen, Quynh Vo, See-Kiong Ng, Bryan Hooi Kuen-Yew
arXiv Computer Vision
Sep 3

TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models

TempoGround is a vision‑language model–native framework for streaming visual grounding that detects cross‑frame object correspondence and explicitly models object presence states. It uses a curriculum prediction mechanism to resolve 2D instance association, predict object entry, continuation, or exit, decode 2D boxes, and lift them to 3D camera‑frame boxes. The approach is further refined with Streaming Grounding Reinforcement, which optimizes grounding, identity, and consistency rewards, and achieves significant improvements on multiple streaming visual grounding benchmarks.

By Leqian Ding, Junning Qiu, Manwen Yang, Yu Guo, Fei Wang
arXiv Machine Learning
Aug 24

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

arXiv:2608.21030v1 Announce Type: cross Abstract: Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottlene...

By Chenghua Zhu, Zhaolu Kang, Qifan Shi, Siyan Wu, Kehan Jiang, Lei Wei, Lianyu Hu, Guangyuan Dong, Mingbo Yang, Rui Lu, Guibo Luo
arXiv Computer Vision
Aug 28

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

PercepCap is a video captioning framework that explicitly models spatio‑temporal perception before generating captions. It follows a perceive‑describe chain, first producing a perception trace of object trajectories and temporal events, then generating the final caption conditioned on that trace. The method uses a two‑stage training strategy—supervised fine‑tuning followed by perception‑grounded reinforcement learning—and builds caption‑aligned perception data to ensure the perception trace and caption refer to the same objects and events.

By Yifan Xu, Zihao Wang, Zhixiao Wang, Jiaming Zhang, Yichun Yang, Desen Meng, Yuanxing Zhang, Pengfei Wan, Limin Wang