arXiv:2608.22064v1 Announce Type: new
Abstract: We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates r...
By Mingqi Gao, Sijie Li, Jungong Han
arXiv:2606. 26734v1 Announce Type: cross Abstract: The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity.
By Priyank Pathak, Mukilan Karuppasamy, Aaditya Baranwal, Shruti Vyas, Yogesh S Rawat
SAM3Dual is a training‑free inference extension of pretrained SAM 3 that won third place in the MOSEv2 track of the 8th Large‑scale Video Object Segmentation Challenge. It separates temporal memory into short‑term and long‑term branches, fuses their responses deterministically, and modulates them with previous‑frame confidence, all while keeping SAM 3 parameters frozen. The approach achieved an official J&F score of 64.37, demonstrating competitive long‑term VOS performance without task‑specific training.
By JeongRae Kim, Chaehyun Kim, Changwon Lim
The paper introduces Event-Driven Refresh + Recurrence Memory (EDRRM) to improve Referring Video Object Segmentation (RVOS). EDRRM selectively re-invokes the Sa2VA model at stable change points, using an EMA‑smoothed event score from tracking cues and a recurrence memory that retrieves anchor frames via CLIP similarity. Experiments on Ref‑DAVIS17, MeViS, and ReVOS show that EDRRM maintains or surpasses J&F scores while reducing refresh calls and false‑positive failures, with modest overhead compared to Sa2VA inference.
By Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo
arXiv:2609.28239v1 Announce Type: new
Abstract: With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulne...
By Li Zeng, Mingcheng Duan, Longfei Fan, Hangtao Zhang, Xianlong Wang, Yanchun Li, Xia Wen, Leo Yu Zhang
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
MoVISA introduces Multi-Token Reasoning for Video Object Segmentation, using multiple segmentation tokens (e.g., SEG0, SEG1) instead of a single token to better localize multiple objects over time. This approach enhances fine-grained alignment between language prompts and spatio-temporal mask predictions, leading to improved performance and interpretability. On benchmarks such as MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS, MoVISA achieves significant gains, notably a 13.2% J and F improvement on MeViS and an 8.4% J and F improvement on ReVOS.
By Ruining Zhao, Ho Kei Cheng, Alexander G Schwing
arXiv:2609.34895v2 Announce Type: replace
Abstract: Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this li...
By Narges Norouzi, Niccol\`{o} Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman, Daan de Geus
GramLoop is a training‑free framework that enhances frozen DINOv3 dense‑prediction models under distribution shift by adding inference computation within the visual backbone. It replays a short transformer window and uses final‑layer cosine‑Gram consistency to control each replay, propagating proposals through the frozen suffix and accepting them via a patchwise gate. Across object detection and semantic segmentation tasks, GramLoop improves performance on all five shifted benchmarks, notably raising COCO‑O mAP by +0.252 and Effective Robustness by +0.250 while maintaining clean ADE20K accuracy.
By Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulnerabilities like backdoor attacks that severely co...
LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.
By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
arXiv:2608. 20107v1 Announce Type: new Abstract: Recent advances in generative video models have significantly improved visual realism in video object removal, yet evaluation protocols still focus on masked region fidelity, treating removal as local inpainting.
By Yigit Ekin, Enes Sanli, Aykut Erdem, Erkut Erdem, Aysegul Dundar