arXiv Computer Vision

When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

The paper introduces FaVOS, a benchmark for Video Object Segmentation (VOS) that focuses on scenarios where target objects appear only intermittently over long videos. It demonstrates that the standard J&F metric can be gamed by empty predictions, allowing trivial models to outperform strong ones like SAM 3. To address this, the authors propose Volumetric J&F, which treats mask sequences as spatio‑temporal volumes, reducing the influence of target‑absence rewards while maintaining sensitivity to segmentation quality and temporal structure.

arXiv Computer Vision
Aug 25

SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge

SAM3Dual is a training‑free inference extension of pretrained SAM 3 that won third place in the MOSEv2 track of the 8th Large‑scale Video Object Segmentation Challenge. It separates temporal memory into short‑term and long‑term branches, fuses their responses deterministically, and modulates them with previous‑frame confidence, all while keeping SAM 3 parameters frozen. The approach achieved an official J&F score of 64.37, demonstrating competitive long‑term VOS performance without task‑specific training.

By JeongRae Kim, Chaehyun Kim, Changwon Lim
arXiv AI
Sep 16

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

VOR-Bench is a new benchmark for video object removal that addresses shortcomings in current evaluation methods by providing a dataset with paired edited videos and graffiti masks, a realistic motion-capable paired-video acquisition framework (rMPAF), and a perception-driven scoring model (VOR-MDSM). The dataset includes diverse data from model-generated, tool-rendered, and camera-captured sources, ensuring robust real-world assessment. Experiments show that VOR-Bench’s evaluation results correlate strongly (ρ > 0.9) with human subjective judgments, bridging the gap between traditional metrics and human preference.

By Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao, Xuchong Zhang, Hongbin Sun, Kongming Liang, Zhanyu Ma
arXiv Machine Learning
Sep 23

SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

SAM‑V is a geometry‑aware extension of the Segment Anything Model (SAM) that integrates 3D priors from a feed‑forward geometry model (VGGT) into 2D segmentation. It uses a prompt‑fusion mechanism to combine sparse SAM prompts with view‑specific camera tokens and local VGGT features, enabling a mask decoder that attends to both dense 2D and 3D cues. The resulting end‑to‑end system produces consistent multi‑view instance segmentation in a single forward pass, achieving significant gains on the IGGT 3D tracking benchmark without offline mask matching or explicit 3D reconstruction.

By Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem
arXiv Computer Vision
3d ago

Toward Open-World Video Segmentation over Long Horizons

The paper introduces Savvy, a zero‑shot, semi‑online, class‑agnostic system that persistently discovers objects and maintains their identities in long videos, and OGA, an evaluation suite that rewards coherent part‑level predictions even when their granularity differs from reference annotations. Savvy combines modular mask discovery, deferred admission based on accumulated evidence, and track consolidation to keep an evolving object set, outperforming DEVA+SAM and EntitySAM on ScanNet and HM3D datasets in metrics such as VPQ_inf, STQ, and AQ. OGA further distinguishes coherent part‑level support from temporal identity failures, revealing that conventional one‑to‑one VPQ metrics are sensitive to annotation granularity and that temporal failures can be detected even when frame‑level masks remain unchanged.

By Qing Su, Kaiyang Li, Yuan Zhuang, Fei Miao, Shihao Ji