arXiv Machine Learning

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.

arXiv AI
Aug 26

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.

By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv Machine Learning
Aug 26

Weakly Supervised Seafloor Segmentation for Seagrass Habitat Mapping in Side-Scan Sonar Imagery

The paper presents a weakly supervised semantic segmentation approach for mapping seagrass habitats using side‑scan sonar imagery. By training a ViT‑based encoder‑decoder with image‑level labels, class activation maps are refined into pseudo‑labels and iteratively self‑trained, achieving high mean intersection‑over‑union scores (up to 89.3 %) without pixel‑level annotations. The method also benefits from self‑supervised pretraining and demonstrates generalizability in field trials.

By Hayat Rajani, Nuno Gracias, Rafael Garcia
arXiv Computer Vision
Aug 26

Comparative Assessment of Deep Learning Architectures for Underwater Subsurface Kelp Forest Segmentation with The Kelp-o-Tron

arXiv:2608.24594v1 Announce Type: new Abstract: Submerged kelp forests are vital coastal ecosystems that support marine biodiversity and ecosystem dynamics, yet accurate underwater kelp segmentation...

By Sundarabalan Balasubramanian, C\'esar Borja, Ana C. Murillo, Lexi N. Wilkes, Meredith L. McPherson, Kira A. Krumhansl, Jennifer A. Dijkstra, Jarrett E. K. Byrnes
arXiv Computer Vision
Aug 31

uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception

uScenes is a new multimodal dataset for underwater robot perception that provides synchronized 3D multibeam sonar point clouds and RGB imagery. It comprises 110 scenes with 95,834 observations, totaling 277.6 minutes of data collected during multiple field sessions. The dataset aims to support research in underwater sensor fusion, cross‑modal representation learning, and 3D scene understanding.

By Trung Tien Dong, Zhenqi Wu, Aditya Penumarti, Zi-Hao Zhang, Micaiah Bartlett, Jane Shin, Xiaomin Lin
arXiv Computer Vision
Sep 14

CoralscapesV2: Panoptic and Fine-Grained Visual Scene Understanding in Coral Reefs

CoralscapesV2 is an expanded dataset for coral reef visual scene understanding, increasing the number of fine‑grained classes from 39 to 95 and adding 65,000 exhaustive fish instance masks. It supports panoptic segmentation by providing high‑quality semantic and instance labels across diverse, unconstrained reef imagery. The dataset serves as a challenging benchmark for modern segmentation models and enables broader applications such as benthic cover mapping and automated fish‑reef interaction analysis.

By Jonathan Sauder, Thomas Ruckli, Gabriel\.e Strodomskyt\.e, Ibrahim Souleiman Abdallah, Rahma Hassan Abdi, Djama Goumaneh Awaleh, Mohamed Houssein Farah, Moustapha Nour, Osama Sharhubil Saad, Mustafa Mohammed Khalafallah Altaib, Maysoon Kteifan, Farah Alsoqi, Eyad Zgool, Jafar Al-Omari, Temesgen Gebremeskel Gebreluel, Zekaria Zekeria Abdulkerim, Meron Ghirmay, Teklehaimanot Beraki, Devis Tuia, Guilhem Banc-Prandi
arXiv Computer Vision
Sep 7

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.

By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
Hugging Face Trending Papers
Jul 27

MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks.

arXiv Computer Vision
Aug 27

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

RSFusionDet introduces a new RGB‑Sonar multimodal object detection dataset (RSFusion) and benchmark metrics for underwater imaging. The proposed detector uses a Cross‑Attention Fusion module to align RGB and sonar features and an Object Matching Head with loss to identify identical objects across modalities. On the RSFusion dataset, RSFusionDet achieves 76.4/48.6 AP for RGB/sonar detection and 83.4 F1‑Score for cross‑modal matching, outperforming existing models and improving over the DINO baseline by 0.7/1.4 AP.

By Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li