arXiv Computer Vision

uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception

uScenes is a new multimodal dataset for underwater robot perception that provides synchronized 3D multibeam sonar point clouds and RGB imagery. It comprises 110 scenes with 95,834 observations, totaling 277.6 minutes of data collected during multiple field sessions. The dataset aims to support research in underwater sensor fusion, cross‑modal representation learning, and 3D scene understanding.

arXiv Computer Vision
Sep 7

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.

By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
arXiv Computer Vision
Aug 27

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

RSFusionDet introduces a new RGB‑Sonar multimodal object detection dataset (RSFusion) and benchmark metrics for underwater imaging. The proposed detector uses a Cross‑Attention Fusion module to align RGB and sonar features and an Object Matching Head with loss to identify identical objects across modalities. On the RSFusion dataset, RSFusionDet achieves 76.4/48.6 AP for RGB/sonar detection and 83.4 F1‑Score for cross‑modal matching, outperforming existing models and improving over the DINO baseline by 0.7/1.4 AP.

By Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li
arXiv Computer Vision
Sep 1

Calibration and Comparative Analysis of Forward-Looking Sonar and 3D Sonar for Enhanced Underwater Object Recognition

arXiv:2608.29433v1 Announce Type: new Abstract: Sonars generate a significant amount of noise. With the advent of new technology capable of producing full 3D point clouds, the noise is amplified in s...

By Aditya Penumarti, Khanh Dong, Zi-Hao Zhang, Yongkyoon Park, Zhenqi Wu, Trung Dong, Shahriar Negahdaripour, Xiaomin Lin, Jane Shin
arXiv AI
Aug 26

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.

By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv Machine Learning
Jul 16

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.

By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
Hugging Face Trending Papers
Jul 23

WAT3R: Feedforward Underwater 3D Reconstruction

Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate multi-view geometry. To address this issue, we propose WAT3R, a feed-forward framework for reconstructing 3D scenes directly from underwater images.

arXiv Computer Vision
Aug 25

Learning Spatially Adaptive Structural Coordination for Underwater Salient Object Detection

The paper introduces SASC-USOD, a framework for underwater salient object detection that learns spatially adaptive coordination between two structural representations: a boundary-sensitive representation using Laplacian filtering and a region-coherent representation via dual-range anisotropic large-kernel aggregation. A spatial coordination module estimates the relative reliability of these representations and adaptively blends them based on image content. Experiments on USOD10K and USOD benchmarks show that SASC-USOD outperforms existing methods, reducing MAE by 4.07% and 23.53% respectively, and its lightweight variant achieves 21 FPS on an NVIDIA Jetson TX2 NX.

By Lin Hong, Chenhui Wang, Linan Deng, Yuning Cui, Yu Zhang, Xin Wang, Bojian Zhang, Xingchen Yang, Fumin Zhang
arXiv Computer Vision
Sep 18

Towards Scaling Marine Perception with Synthetic Data

The paper introduces an extension to the OceanSim underwater perception simulator, adding a Synthetic Data Generation pipeline that produces large, automatically labeled, photorealistic datasets with configurable scene and sensor settings. The authors evaluate this pipeline on a real-world sea urchin detection task, examining how different synthetic scene variations influence sim-to-real performance. They discuss the pipeline’s findings, limitations, and future directions for improving rendering fidelity, scene diversity, and sim-to-real generalization.

By Haoyu Ma, Onur Bagoren, Anja Sheppard, Elias Fandi, Ashrith Edukulla, Tanner Aslan, Natasha Sieh, Jingyu Song, Katherine A. Skinner