arXiv:2609.39785v1 Announce Type: new
Abstract: The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised...
By Weijian Jian, Xiaoyue Zhang, Bin Xiao, Chunyu Xie, Yixiao He, Yutao Liu, Dawei Leng, Yuhui Yin
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv:2603. 13994v2 Announce Type: replace-cross Abstract: Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties.
By Hossein Adeli, Seoyoung Ahn, Andrew Luo, Mengmi Zhang, Nikolaus Kriegeskorte, Gregory Zelinsky
arXiv:2509.22650v3 Announce Type: replace
Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
By Anna Kukleva, Enis Simsar, Alessio Tonioni, Muhammad Ferjad Naeem, Federico Tombari, Jan Eric Lenssen, Bernt Schiele
arXiv:2606. 14765v1 Announce Type: cross Abstract: Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning.
By Qinwu Xu
arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.
By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision.
arXiv:2606. 06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding.
By Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen
DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.
By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.
The paper introduces MiRA, a plug‑in framework that reweights framewise attention in Vision Transformer video models to better capture subtle facial dynamics for expression recognition. MiRA computes frame‑level confidence and intra‑frame concentration from self‑attention maps, redistributing attention toward localized facial cues without adding trainable parameters. Two modes—an exact post‑softmax redistribution and a lightweight flashLite pre‑softmax approximation—are proposed, and experiments on facial expression recognition benchmarks show consistent gains over strong ViT baselines.
By Seongro Yoon, Donghyeon Cho, Jinsun Park, Fran\c{c}ois Br\'emond
arXiv:2605.12491v2 Announce Type: replace
Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...
By Alan Z. Song, Yinjie Chen, Mu Nan, Deva Ramanan, Michael J. Tarr, Andrew F. Luo