arXiv Computer Vision

GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring

Hugging Face Trending Papers
Aug 5

Promptable Animal Pose Tracking Across Species

Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data.

arXiv Computer Vision
Sep 7

PuTR-CouT: Counting-by-Tracking in Camera-Trap Image Sequences

PuTR-CouT is a transformer‑based counting‑by‑tracking framework designed for camera‑trap image sequences. It generates synthetic training data using structural priors to create pseudo‑tracking labels, enabling the tracker to associate detections across frames and estimate per‑species counts. The method improves upon the MaxBoxCount baseline on the iWildCam 2021 benchmark, offering competitive counting results along with multi‑species predictions and track‑level verification.

By Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos
arXiv Computer Vision
Sep 4

Counting Animals in Camera-Traps Image Sequences without Count Labels: Winning Solution to the iWildCam 2021 Challenge

The paper presents MaxBoxCount, the winning solution to the iWildCam 2021 Challenge, which tackles counting animals in camera‑trap image sequences without using count labels. It combines a robust species classification pipeline with a counting heuristic based on MegaDetector detections to estimate the number of unique individuals across short image bursts. The method addresses challenges posed by temporal discontinuities and the high cost of manual count annotations.

By Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos
arXiv AI
Sep 10

What Does Animal Re-Identification Learn? Linear Biological Concepts and Their Origins in Visual Representations

The study investigates whether Vision Transformer (ViT)-based animal re-identification models learn biologically meaningful concepts. Using a DINOv3 backbone fine‑tuned on Western lowland gorilla images, the authors find that sex and age emerge as linear directions in the model’s representations, generalizing to unseen individuals with high AUROC scores. They demonstrate that the sex direction is causally used by the model, that fine‑tuning relocates these concepts within the network, and that the representations reflect a graded biological axis encoded redundantly across the population.

By Robert Nolting, Alexandra Schild, Moritz Weckbecker, Maximilian Schall, Gerard de Melo
arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv AI
Jun 2

CAFOSat: A Strongly Annotated Dataset for Infrastructure-Aware CAFO Mapping Using High-Resolution Imagery

arXiv:2606. 00548v1 Announce Type: cross Abstract: Concentrated Animal Feeding Operations (CAFOs) play an important role in agricultural production but are also associated with environmental, public health, and disease surveillance concerns.

By Oishee Bintey Hoque, Nibir Chandra Mandal, Mandy L Wilson, Samarth Swarup, Madhav Marathe, Abhijin Adiga