arXiv AI By Tianshu Zhang, Yan Wang, Ji Qi, Lijie Wen

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

Read the original on arXiv AI →

arXiv:2606. 29023v1 Announce Type: cross Abstract: Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
5d ago

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

arXiv:2608. 13344v1 Announce Type: new Abstract: Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences.

By Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang