Hugging Face Trending Papers

Whareformer: Learning to Track What is Where in Long Egocentric Videos

Read the original on Hugging Face Trending Papers →

The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become heavily occluded. In this paper, we propose the first learning-based solution to the OSNOM task: Whareformer, a transformer-based model with two components: an updatable memory of established tracks and a track assignment module that associates observations with existing tracks in a feed-forward manner.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 25

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

TrackEverything is a 3D point tracker that overcomes the trade‑off between sparse long‑horizon tracking and dense short‑clip tracking by representing videos as persistent 3D scene tracks in world coordinates. It introduces voxel‑based de‑duplication at sliding‑window boundaries, a two‑stage refinement process (endpoint refiner and lightweight trajectory refiner), and a 3D WAFT module that replaces memory‑heavy 4D correlation volumes with efficient feature sampling. The method can track all visible points in videos longer than 1000 frames using only 40 GB of GPU memory, outperforming existing dense trackers on short clips and matching sparse trackers on long sequences.

By Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley
arXiv Computer Vision
3d ago

DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention

DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.

By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung