iSEE: Object Permanence Through Self-Supervision
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become heavily occluded. In this paper, we propose the first learning-based solution to the OSNOM task: Whareformer, a transformer-based model with two components: an updatable memory of established tracks and a track assignment module that associates observations with existing tracks in a feed-forward manner.
arXiv:2608.28216v1 Announce Type: new Abstract: Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is ab...
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
The paper introduces S$^3$T, a fully self‑contained framework for continuous video state tracking that uses temporal self‑distillation. It treats denser temporal sampling as privileged information, letting a dense‑view teacher guide a sparse‑view student to match its next‑token distribution without external labels or reward signals. Experiments on LLaVA-OneVision-2-8B show significant accuracy gains on VSTAT and MVBench benchmarks, and the learned capability transfers from synthetic to real videos.
arXiv:2602. 14771v5 Announce Type: replace-cross Abstract: The human visual system tracks objects by integrating current observations with previously observed information, adapting to target and scene changes, and reasoning about occlusion at fine granularity.
arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.