The paper presents a method that leverages sparse expert point annotations from historical benthic surveys to improve dense segmentation of marine imagery. By using these points as visual prompts for the SAM2 foundation model and introducing a mechanism to filter out unreliable points, the authors generate high‑quality pseudo‑ground‑truth masks that train more accurate fine‑grained semantic segmentation models. The approach is validated on public benthic datasets and a new benchmark featuring real‑world sparse annotations, aiming to enable scalable ecological analysis.
By Cesar Borja, Breck A. McCollum, Jarret E. Byrnes, Kenneth Sebens, Ana C. Murillo
arXiv:2607. 24064v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs.
By Tuan-An To, Yuk-Kwan Wong, Tuan-Anh Vu, Ziqiang Zheng, Sai-Kit Yeung
CoralscapesV2 is an expanded dataset for coral reef visual scene understanding, increasing the number of fine‑grained classes from 39 to 95 and adding 65,000 exhaustive fish instance masks. It supports panoptic segmentation by providing high‑quality semantic and instance labels across diverse, unconstrained reef imagery. The dataset serves as a challenging benchmark for modern segmentation models and enables broader applications such as benthic cover mapping and automated fish‑reef interaction analysis.
By Jonathan Sauder, Thomas Ruckli, Gabriel\.e Strodomskyt\.e, Ibrahim Souleiman Abdallah, Rahma Hassan Abdi, Djama Goumaneh Awaleh, Mohamed Houssein Farah, Moustapha Nour, Osama Sharhubil Saad, Mustafa Mohammed Khalafallah Altaib, Maysoon Kteifan, Farah Alsoqi, Eyad Zgool, Jafar Al-Omari, Temesgen Gebremeskel Gebreluel, Zekaria Zekeria Abdulkerim, Meron Ghirmay, Teklehaimanot Beraki, Devis Tuia, Guilhem Banc-Prandi
The paper introduces Molmo2Fish, an interactive tool that uses a multimodal large language model to correct imperfect fish tracking predictions in sonar datasets. It demonstrates that the system can effectively improve multi‑object tracking performance through natural language guidance, though further enhancements are needed. The authors provide open‑source code and data for replication.
PuTR-CouT is a transformer‑based counting‑by‑tracking framework designed for camera‑trap image sequences. It generates synthetic training data using structural priors to create pseudo‑tracking labels, enabling the tracker to associate detections across frames and estimate per‑species counts. The method improves upon the MaxBoxCount baseline on the iWildCam 2021 benchmark, offering competitive counting results along with multi‑species predictions and track‑level verification.
By Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos
arXiv:2607. 09876v1 Announce Type: cross Abstract: Automatically retrieving videos from large camera-trap datasets remains challenging.
By Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, Devis Tuia, Alexander Mathis
arXiv:2606. 00080v1 Announce Type: cross Abstract: Marine plankton underpin aquatic food webs and play a key role in global CO2 sequestration, making reliable species identification critical for understanding ocean health and climate feedbacks.
By Alan Gerson Contreras Montanares, Luis Valenzuela, Luis Mart\'i, Nayat Sanchez-Pi
arXiv:2512.07776v2 Announce Type: replace
Abstract: Monitoring critically endangered western lowland gorillas is currently hampered by the immense manual effort required to re-identify individuals fr...
By Maximilian Schall, Felix Leonard Kn\"ofel, Noah Elias K\"onig, Jan Jonas Kubeler, Maximilian von Klinski, Joan Wilhelm Linnemann, Xiaoshi Liu, Iven Jelle Schlegelmilch, Ole Woyciniuk, Alexandra Schild, Dante Wasmuht, Magdalena Bermejo Espinet, German Illera Basas, Gerard de Melo
arXiv:2607. 09443v1 Announce Type: cross Abstract: Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts.
By Anil Osman Tur, Tonje Knutsen Sordalen, Kim Tallaksen Halvorsen, Cigdem Beyan
The paper presents MaxBoxCount, the winning solution to the iWildCam 2021 Challenge, which tackles counting animals in camera‑trap image sequences without using count labels. It combines a robust species classification pipeline with a counting heuristic based on MegaDetector detections to estimate the number of unique individuals across short image bursts. The method addresses challenges posed by temporal discontinuities and the high cost of manual count annotations.
By Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos
Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward vi...
Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data.