WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper presents a method that leverages sparse expert point annotations from historical benthic surveys to improve dense segmentation of marine imagery. By using these points as visual prompts for the SAM2 foundation model and introducing a mechanism to filter out unreliable points, the authors generate high‑quality pseudo‑ground‑truth masks that train more accurate fine‑grained semantic segmentation models. The approach is validated on public benthic datasets and a new benchmark featuring real‑world sparse annotations, aiming to enable scalable ecological analysis.
arXiv:2607. 24064v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs.
CoralscapesV2 is an expanded dataset for coral reef visual scene understanding, increasing the number of fine‑grained classes from 39 to 95 and adding 65,000 exhaustive fish instance masks. It supports panoptic segmentation by providing high‑quality semantic and instance labels across diverse, unconstrained reef imagery. The dataset serves as a challenging benchmark for modern segmentation models and enables broader applications such as benthic cover mapping and automated fish‑reef interaction analysis.
The paper introduces Molmo2Fish, an interactive tool that uses a multimodal large language model to correct imperfect fish tracking predictions in sonar datasets. It demonstrates that the system can effectively improve multi‑object tracking performance through natural language guidance, though further enhancements are needed. The authors provide open‑source code and data for replication.
PuTR-CouT is a transformer‑based counting‑by‑tracking framework designed for camera‑trap image sequences. It generates synthetic training data using structural priors to create pseudo‑tracking labels, enabling the tracker to associate detections across frames and estimate per‑species counts. The method improves upon the MaxBoxCount baseline on the iWildCam 2021 benchmark, offering competitive counting results along with multi‑species predictions and track‑level verification.
arXiv:2607. 09876v1 Announce Type: cross Abstract: Automatically retrieving videos from large camera-trap datasets remains challenging.