arXiv Computer Vision

Toward Open-World Video Segmentation over Long Horizons

The paper introduces Savvy, a zero‑shot, semi‑online, class‑agnostic system that persistently discovers objects and maintains their identities in long videos, and OGA, an evaluation suite that rewards coherent part‑level predictions even when their granularity differs from reference annotations. Savvy combines modular mask discovery, deferred admission based on accumulated evidence, and track consolidation to keep an evolving object set, outperforming DEVA+SAM and EntitySAM on ScanNet and HM3D datasets in metrics such as VPQ_inf, STQ, and AQ. OGA further distinguishes coherent part‑level support from temporal identity failures, revealing that conventional one‑to‑one VPQ metrics are sensitive to annotation granularity and that temporal failures can be detected even when frame‑level masks remain unchanged.

arXiv Computer Vision
Sep 24

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.

By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
arXiv Machine Learning
Sep 23

SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.

By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem
arXiv Computer Vision
5d ago

When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

The paper introduces FaVOS, a benchmark for Video Object Segmentation (VOS) that focuses on scenarios where target objects appear only intermittently over long videos. It demonstrates that the standard J&F metric can be gamed by empty predictions, allowing trivial models to outperform strong ones like SAM 3. To address this, the authors propose Volumetric J&F, which treats mask sequences as spatio‑temporal volumes, reducing the influence of target‑absence rewards while maintaining sensitivity to segmentation quality and temporal structure.

By Jihwan Hong, Woohyeon Park, Jaeik Kim, Jaeyoung Do
arXiv AI
Jun 11

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.

By Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan
arXiv Computer Vision
Oct 1

Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation

The paper introduces Event-Driven Refresh + Recurrence Memory (EDRRM) to improve Referring Video Object Segmentation (RVOS). EDRRM selectively re-invokes the Sa2VA model at stable change points, using an EMA‑smoothed event score from tracking cues and a recurrence memory that retrieves anchor frames via CLIP similarity. Experiments on Ref‑DAVIS17, MeViS, and ReVOS show that EDRRM maintains or surpasses J&F scores while reducing refresh calls and false‑positive failures, with modest overhead compared to Sa2VA inference.

By Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo
arXiv AI
Sep 25

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.

By Jerrin Bright, John Zelek