arXiv Computer Vision
Aug 25

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

arXiv:2501.04001v4 Announce Type: replace Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...

By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv Computer Vision
Sep 4

ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking

The paper introduces ORMOT, a new task that extends Referring Multi‑Object Tracking to omnidirectional 360° imagery, ensuring full scene context for language‑guided tracking. It presents ORSet, a dataset of 27 omnidirectional scenes with 848 language descriptions and 3,401 annotated objects, and introduces ORTrack, an LVLM‑driven framework that performs zero‑shot detection and robust cross‑frame association. Experiments on ORSet show that ORTrack achieves state‑of‑the‑art performance, establishing a strong baseline for future research.

By Zihan Zhou, Sijia Chen, Yanqiu Yu, En Yu, Wenbing Tao