arXiv Computer Vision

TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos

arXiv Computer Vision
1d ago

Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

arXiv:2609.36217v1 Announce Type: new Abstract: A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from vi...

By Xinming Dai, Qihang Jin, Tianshu Tan, Baiyuan Chen, Hanrui Lyu, Lenny Aharon, Kyle Daruwalla, Xun Helen Hou, Matthew R. Whiteway, Liam Paninski, Yizi Zhang
arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv Computer Vision
Aug 24

WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition

arXiv:2608.21281v1 Announce Type: new Abstract: Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging...

By Abigail G. Grassick, Jerome Tze-Hou Hsu, Ethan Lin, Ziang Liu, Max Whitton, Madelyn Hair, Liam Gutierrez, Haozheng Yu, Kristin Branson, Vivek Jayaraman, Michael A. Gil, Andrew M. Hein, Jennifer J. Sun
Hugging Face Trending Papers
3d ago

VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

VastMAT is a large‑scale multi‑animal tracking benchmark featuring 2,947 videos, 337 animal categories, and over 3.6 million bounding boxes with 22,883 identity trajectories. It emphasizes high‑quality, expert‑reviewed annotations and introduces Seen‑category and Unseen‑category evaluation protocols, revealing significant challenges in tracking unseen animals. The authors also propose a lightweight Center‑Distance‑Augmented Association module that boosts HOTA scores for existing MOT methods without extra training.

arXiv AI
6d ago

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

TrackEverything is a 3D point tracker that overcomes the trade‑off between sparse long‑horizon tracking and dense short‑clip tracking by representing videos as persistent 3D scene tracks in world coordinates. It introduces voxel‑based de‑duplication at sliding‑window boundaries, a two‑stage refinement process (endpoint refiner and lightweight trajectory refiner), and a 3D WAFT module that replaces memory‑heavy 4D correlation volumes with efficient feature sampling. The method can track all visible points in videos longer than 1000 frames using only 40 GB of GPU memory, outperforming existing dense trackers on short clips and matching sparse trackers on long sequences.

By Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley
arXiv AI
23h ago

Emergent Multi-View Geometry Through Self-Distillation

The paper introduces Poincar3, a self‑supervised method that learns multi‑view representations through self‑distillation rather than RGB reconstruction. By combining masked patch and image‑level distillation with a teacher that sees additional views, it trains from scratch without explicit 3D supervision. Poincar3 surpasses prior single‑ and multi‑view self‑supervised methods on tasks such as correspondence estimation, camera pose estimation, and 3D reconstruction, and its features encode camera motion more accurately thanks to a lightweight Poincaré adapter.

By David Nordstr\"om, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl