CoralscapesV2 is an expanded dataset for coral reef visual scene understanding, increasing the number of fine‑grained classes from 39 to 95 and adding 65,000 exhaustive fish instance masks. It supports panoptic segmentation by providing high‑quality semantic and instance labels across diverse, unconstrained reef imagery. The dataset serves as a challenging benchmark for modern segmentation models and enables broader applications such as benthic cover mapping and automated fish‑reef interaction analysis.
By Jonathan Sauder, Thomas Ruckli, Gabriel\.e Strodomskyt\.e, Ibrahim Souleiman Abdallah, Rahma Hassan Abdi, Djama Goumaneh Awaleh, Mohamed Houssein Farah, Moustapha Nour, Osama Sharhubil Saad, Mustafa Mohammed Khalafallah Altaib, Maysoon Kteifan, Farah Alsoqi, Eyad Zgool, Jafar Al-Omari, Temesgen Gebremeskel Gebreluel, Zekaria Zekeria Abdulkerim, Meron Ghirmay, Teklehaimanot Beraki, Devis Tuia, Guilhem Banc-Prandi
arXiv:2608.21281v1 Announce Type: new
Abstract: Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging...
By Abigail G. Grassick, Jerome Tze-Hou Hsu, Ethan Lin, Ziang Liu, Max Whitton, Madelyn Hair, Liam Gutierrez, Haozheng Yu, Kristin Branson, Vivek Jayaraman, Michael A. Gil, Andrew M. Hein, Jennifer J. Sun
arXiv:2609.23397v1 Announce Type: new
Abstract: Shrimp diseases continue to cause devastating losses in the aquaculture industry, driving a critical need for robust, automated detection. This work co...
By Vinh Canh-Thanh Truong, Hai-Binh Pham, Ngoc Hong Tran
arXiv:2609.38347v1 Announce Type: new
Abstract: Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visuall...
By Patt Phurtivilai, Zhiyang Dou, Yifan Wu, Kinfung Chu, Yuan Liu, Lei Yang, Wenping Wang, Taku Komura
arXiv:2609.25500v1 Announce Type: new
Abstract: Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. When using object detection mod...
By Lonny Lundsten, Kevin Barnard, Dave Caress
The paper introduces Distortion Extenders (DEX), learnable parameters that adapt vision foundation models to fisheye cameras by modeling distortion coefficients and correcting distributional shifts between fisheye and perspective images. DEX is applied to monocular depth estimation and open‑vocabulary segmentation across convolutional and Transformer architectures, consistently outperforming baselines on indoor and outdoor fisheye datasets. Additionally, DEX activations can be decoded to obtain distortion coefficients, aiding camera calibration.
By Rit Gangopadhyay, Alex Wong
DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.
By Quang-Trung Truong, Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
The paper introduces Molmo2Fish, an interactive tool that uses a multimodal large language model to correct imperfect fish tracking predictions in sonar datasets. It demonstrates that the system can effectively improve multi‑object tracking performance through natural language guidance, though further enhancements are needed. The authors provide open‑source code and data for replication.
arXiv:2609.15484v1 Announce Type: new
Abstract: We report on the continued development of CatchMonitor, resulting in a prototype computer vision system designed to automatically quantify discarded fi...
By Geoff French, Michal Mackiewicz, Mark Fisher, Helen Holah, Rebecca Lamb
arXiv:2509.06422v2 Announce Type: replace
Abstract: Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods...
By Hua Zhang, Changjiang Luo
The paper evaluates the new YOLO26 architecture, which offers NMS-free end-to-end inference and is tailored for CPU-based edge devices, against three earlier Ultralytics models (YOLOv5u, YOLOv8, and YOLO11) in aquaculture fish mortality detection. Across nano, small, and medium scales, all models achieved similar detection accuracy on a full dataset, but differences emerged in data efficiency and deployment performance: YOLOv8 reached 90% mAP50 with only 400 images, while YOLO26 variants needed 1,000 images; YOLO26n was fastest on a Raspberry Pi 5 (7.51 FPS), whereas YOLOv5mu led on CPU-based hardware. The study concludes that architectural novelty alone does not dictate suitability for edge AI in aquaculture; training data size, target hardware, and inference needs must be jointly considered.
By Rakesh Ranjan, Gajanan S. Kothawade, Kata Sharrer, Scott Tsukuda, Christopher Good
arXiv:2608. 11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming.
By Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen