uScenes is a new multimodal dataset for underwater robot perception that provides synchronized 3D multibeam sonar point clouds and RGB imagery. It comprises 110 scenes with 95,834 observations, totaling 277.6 minutes of data collected during multiple field sessions. The dataset aims to support research in underwater sensor fusion, cross‑modal representation learning, and 3D scene understanding.
By Trung Tien Dong, Zhenqi Wu, Aditya Penumarti, Zi-Hao Zhang, Micaiah Bartlett, Jane Shin, Xiaomin Lin
AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.
By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
arXiv:2608.23479v1 Announce Type: new
Abstract: Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlen...
By Taqi Hamoda, Nuno Gracias
Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings.
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate multi-view geometry. To address this issue, we propose WAT3R, a feed-forward framework for reconstructing 3D scenes directly from underwater images.
arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.
By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.
By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv:2609.25271v1 Announce Type: cross
Abstract: Sidescan sonar is a common sensor for both manned and autonomous marine exploration and mapping, yet very few methods build upon or exploit the geome...
By Kalin Norman, Joshua G. Mangelson
arXiv:2608.22906v1 Announce Type: new
Abstract: Recent monocular 3D Gaussian Splatting (3DGS) streaming reconstruction methods have achieved impressive performance by balancing reconstruction quality...
By Yingxiang Xu, Kerui Ren, Wenqi Guo, Changjian Jiang, Tao Lu, Linning Xu, Mulin Yu
OceanXL is a new framework that applies 3D Gaussian Splatting to large-scale underwater scenes by partitioning them into spatially coherent blocks and using adaptive pruning to remove redundant primitives. This divide‑and‑conquer approach improves training efficiency and rendering performance while maintaining global geometric consistency. The authors also release a large underwater dataset and demonstrate that OceanXL achieves favorable scalability, compactness, and efficiency compared to existing baselines, with competitive quality on smaller datasets and smaller model sizes than other underwater methods.
By Haoran Wang, Shaoyu Cai, Adrian Azzarelli, Zhuodong Jiang, Guoxi Huang, Eng Tat Khoo, Brett Seymour, Fan Zhang, David Bull, Nantheera Anantrasirichai
The paper introduces an extension to the OceanSim underwater perception simulator, adding a Synthetic Data Generation pipeline that produces large, automatically labeled, photorealistic datasets with configurable scene and sensor settings. The authors evaluate this pipeline on a real-world sea urchin detection task, examining how different synthetic scene variations influence sim-to-real performance. They discuss the pipeline’s findings, limitations, and future directions for improving rendering fidelity, scene diversity, and sim-to-real generalization.
By Haoyu Ma, Onur Bagoren, Anja Sheppard, Elias Fandi, Ashrith Edukulla, Tanner Aslan, Natasha Sieh, Jingyu Song, Katherine A. Skinner
The paper presents a framework that builds a static point cloud prior map from past camera traversals, augmenting each point with DINOv3 semantic features. During runtime, a local prior patch is retrieved, encoded with a sparse voxel backbone, and fused with lifted multi‑view camera features in bird’s‑eye view. This fused representation is then used by sparse transformer heads to predict 3D objects and vectorized map elements, achieving improved performance on Argoverse 2 without requiring LiDAR for prior‑map construction or online inference.
By Markus K\"appeler, Rohit Mohan, Abhinav Valada