arXiv Computer Vision

Robust Promptable Video Object Segmentation

arXiv Computer Vision
Aug 25

SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge

SAM3Dual is a training‑free inference extension of pretrained SAM 3 that won third place in the MOSEv2 track of the 8th Large‑scale Video Object Segmentation Challenge. It separates temporal memory into short‑term and long‑term branches, fuses their responses deterministically, and modulates them with previous‑frame confidence, all while keeping SAM 3 parameters frozen. The approach achieved an official J&F score of 64.37, demonstrating competitive long‑term VOS performance without task‑specific training.

By JeongRae Kim, Chaehyun Kim, Changwon Lim
arXiv Computer Vision
1d ago

Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation

The paper introduces Event-Driven Refresh + Recurrence Memory (EDRRM) to improve Referring Video Object Segmentation (RVOS). EDRRM selectively re-invokes the Sa2VA model at stable change points, using an EMA‑smoothed event score from tracking cues and a recurrence memory that retrieves anchor frames via CLIP similarity. Experiments on Ref‑DAVIS17, MeViS, and ReVOS show that EDRRM maintains or surpasses J&F scores while reducing refresh calls and false‑positive failures, with modest overhead compared to Sa2VA inference.

By Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo
arXiv Computer Vision
Sep 25

MoVISA: Multi-Token Reasoning for Video Object Segmentation

MoVISA introduces Multi-Token Reasoning for Video Object Segmentation, using multiple segmentation tokens (e.g., SEG0, SEG1) instead of a single token to better localize multiple objects over time. This approach enhances fine-grained alignment between language prompts and spatio-temporal mask predictions, leading to improved performance and interpretability. On benchmarks such as MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS, MoVISA achieves significant gains, notably a 13.2% J and F improvement on MeViS and an 8.4% J and F improvement on ReVOS.

By Ruining Zhao, Ho Kei Cheng, Alexander G Schwing
arXiv Computer Vision
Sep 1

GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction

GramLoop is a training‑free framework that enhances frozen DINOv3 dense‑prediction models under distribution shift by adding inference computation within the visual backbone. It replays a short transformer window and uses final‑layer cosine‑Gram consistency to control each replay, propagating proposals through the frozen suffix and accepting them via a patchwise gate. Across object detection and semantic segmentation tasks, GramLoop improves performance on all five shifted benchmarks, notably raising COCO‑O mAP by +0.252 and Effective Robustness by +0.250 while maintaining clean ADE20K accuracy.

By Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
arXiv Computer Vision
Sep 24

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.

By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni