arXiv Computer Vision By Qing Su, Kaiyang Li, Yuan Zhuang, Fei Miao, Shihao Ji

Toward Open-World Video Segmentation over Long Horizons

Read the original on arXiv Computer Vision →

The paper introduces Savvy, a zero‑shot, semi‑online, class‑agnostic system that persistently discovers objects and maintains their identities in long videos, and OGA, an evaluation suite that rewards coherent part‑level predictions even when their granularity differs from reference annotations. Savvy combines modular mask discovery, deferred admission based on accumulated evidence, and track consolidation to keep an evolving object set, outperforming DEVA+SAM and EntitySAM on ScanNet and HM3D datasets in metrics such as VPQ_inf, STQ, and AQ. OGA further distinguishes coherent part‑level support from temporal identity failures, revealing that conventional one‑to‑one VPQ metrics are sensitive to annotation granularity and that temporal failures can be detected even when frame‑level masks remain unchanged.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 24

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.

By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
arXiv Machine Learning
Sep 23

SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.

By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem