arXiv:2608.22064v1 Announce Type: new
Abstract: We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates r...
By Mingqi Gao, Sijie Li, Jungong Han
arXiv:2509.22650v3 Announce Type: replace
Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
By Anna Kukleva, Enis Simsar, Alessio Tonioni, Muhammad Ferjad Naeem, Federico Tombari, Jan Eric Lenssen, Bernt Schiele
arXiv:2603. 24016v2 Announce Type: replace-cross Abstract: Multi-Object Tracking (MOT) has traditionally focused on a few specific categories, restricting its applicability to real-world scenarios involving diverse objects.
By Zekun Qian, Wei Feng, Ruize Han, Junhui Hou
arXiv:2501.04001v4 Announce Type: replace
Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...
By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
The paper introduces ORMOT, a new task that extends Referring Multi‑Object Tracking to omnidirectional 360° imagery, ensuring full scene context for language‑guided tracking. It presents ORSet, a dataset of 27 omnidirectional scenes with 848 language descriptions and 3,401 annotated objects, and introduces ORTrack, an LVLM‑driven framework that performs zero‑shot detection and robust cross‑frame association. Experiments on ORSet show that ORTrack achieves state‑of‑the‑art performance, establishing a strong baseline for future research.
By Zihan Zhou, Sijia Chen, Yanqiu Yu, En Yu, Wenbing Tao
arXiv:2609.37042v1 Announce Type: cross
Abstract: Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the lar...
By Shuo Yang, Changbai Li, Rui Tang, Xinyu Zhao, Linlin Yang, Baochang Zhang
arXiv:2609.39785v1 Announce Type: new
Abstract: The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised...
By Weijian Jian, Xiaoyue Zhang, Bin Xiao, Chunyu Xie, Yixiao He, Yutao Liu, Dawei Leng, Yuhui Yin
arXiv:2609.09396v1 Announce Type: new
Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...
By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
The paper introduces Event-Driven Refresh + Recurrence Memory (EDRRM) to improve Referring Video Object Segmentation (RVOS). EDRRM selectively re-invokes the Sa2VA model at stable change points, using an EMA‑smoothed event score from tracking cues and a recurrence memory that retrieves anchor frames via CLIP similarity. Experiments on Ref‑DAVIS17, MeViS, and ReVOS show that EDRRM maintains or surpasses J&F scores while reducing refresh calls and false‑positive failures, with modest overhead compared to Sa2VA inference.
By Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo
arXiv:2608.29126v1 Announce Type: new
Abstract: Referring single-object tracking enables language-grounded target initialization and subsequent tracking by jointly leveraging semantic cues and visual...
By Han Wang, Yuxuan Liu, Yuhan Sun, Jian Yang, Xiaotong Xu, Yixuan Lv, Zhuang Zhou, Shengyang Li
arXiv:2609.37339v1 Announce Type: new
Abstract: General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on boundin...
By Jer Pelhan, Alan Lukezic, Matej Kristan
Segmentation-based multi-object tracking (MOT) with foundation video models such as SAM2 offers strong localization quality, yet remains fragile in crowded, real-world scenes. In detector-prompted SAM...