uScenes is a new multimodal dataset for underwater robot perception that provides synchronized 3D multibeam sonar point clouds and RGB imagery. It comprises 110 scenes with 95,834 observations, totaling 277.6 minutes of data collected during multiple field sessions. The dataset aims to support research in underwater sensor fusion, cross‑modal representation learning, and 3D scene understanding.
By Trung Tien Dong, Zhenqi Wu, Aditya Penumarti, Zi-Hao Zhang, Micaiah Bartlett, Jane Shin, Xiaomin Lin
arXiv:2608.29433v1 Announce Type: new
Abstract: Sonars generate a significant amount of noise. With the advent of new technology capable of producing full 3D point clouds, the noise is amplified in s...
By Aditya Penumarti, Khanh Dong, Zi-Hao Zhang, Yongkyoon Park, Zhenqi Wu, Trung Dong, Shahriar Negahdaripour, Xiaomin Lin, Jane Shin
Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings.
arXiv:2608. 19710v1 Announce Type: cross Abstract: Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur.
By Mohammad Arif Ul Alam
arXiv:2608.23479v1 Announce Type: new
Abstract: Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlen...
By Taqi Hamoda, Nuno Gracias
arXiv:2608.22906v1 Announce Type: new
Abstract: Recent monocular 3D Gaussian Splatting (3DGS) streaming reconstruction methods have achieved impressive performance by balancing reconstruction quality...
By Yingxiang Xu, Kerui Ren, Wenqi Guo, Changjian Jiang, Tao Lu, Linning Xu, Mulin Yu
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate multi-view geometry. To address this issue, we propose WAT3R, a feed-forward framework for reconstructing 3D scenes directly from underwater images.
Recent monocular 3D Gaussian Splatting (3DGS) streaming reconstruction methods have achieved impressive performance by balancing reconstruction quality and efficiency. However, extending these framewo...
SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.
By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
RSFusionDet introduces a new RGB‑Sonar multimodal object detection dataset (RSFusion) and benchmark metrics for underwater imaging. The proposed detector uses a Cross‑Attention Fusion module to align RGB and sonar features and an Object Matching Head with loss to identify identical objects across modalities. On the RSFusion dataset, RSFusionDet achieves 76.4/48.6 AP for RGB/sonar detection and 83.4 F1‑Score for cross‑modal matching, outperforming existing models and improving over the DINO baseline by 0.7/1.4 AP.
By Zhuoyan Liu, Yihan Wang, Bo Wang, Bing Wang, Ye Li
arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.
By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
By Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal