AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.
By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
arXiv:2608.23479v1 Announce Type: new
Abstract: Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlen...
By Taqi Hamoda, Nuno Gracias
arXiv:2608. 08025v1 Announce Type: cross Abstract: Accurate 3D reconstruction of ships at sea is important for maritime supervision, damage assessment, and autonomous maritime operations.
By Jiaming Chen, Juntao Yang, Zhentao Zou, Qi Ming, Yi Yu, Zhihang Zhong, Xue Yang, Xue Jiang, Yue Zhou
arXiv:2608.21276v1 Announce Type: cross
Abstract: Coastal environments contain rich, largely unexploited geometric structure capable of providing globally referenced localization cues. In this work,...
By Derek R. Benham, Joshua G. Mangelson
arXiv:2608.29433v1 Announce Type: new
Abstract: Sonars generate a significant amount of noise. With the advent of new technology capable of producing full 3D point clouds, the noise is amplified in s...
By Aditya Penumarti, Khanh Dong, Zi-Hao Zhang, Yongkyoon Park, Zhenqi Wu, Trung Dong, Shahriar Negahdaripour, Xiaomin Lin, Jane Shin
arXiv:2506. 22174v3 Announce Type: replace-cross Abstract: The transport industry has recently shown significant interest in unmanned surface vehicles (USVs), specifically for port and inland waterway transport.
By Bavo Lesy, Siemen Herremans, Robin Kerstens, Jan Steckel, Walter Daems, Siegfried Mercelis, Ali Anwar
arXiv:2608.29315v1 Announce Type: cross
Abstract: This work introduces Semantically-Guided Exploration (SGE), a modular exploration framework for ground vehicles that integrates pixel-level semantic...
By Christopher Tatsch, Yu Gu
ALS boresight calibration has relied for two decades on dedicated flight patterns over structured scenes containing planar surfaces of varied aspect and slope. While reliable, this approach imposes constraints on the scene content and operations, which limits its applicability to boresight recovery within routine mapping missions.
RiverVLN introduces the first benchmark for long‑horizon vision‑language navigation (VLN) of unmanned surface vehicles (USVs) in continuous riverine motion. The PGT‑NAV framework converts navigation instructions into an ordered sequence of visually verifiable semantic phases, maintaining an active phase online through grounded visual and motion evidence. This phase‑grounded approach reduces recursive position and heading drift, achieving a 0.79 success rate in Unity‑ROS closed‑loop tests and demonstrating transfer to real‑world USV deployment.
By Jieling Wu, Yuehao Huang, Jiajun Lv, Tao Huang, Yong Liu, Weiwei Liu
arXiv:2608.30471v1 Announce Type: new
Abstract: This paper investigates the problem of position estimation of unmanned surface vessels (USVs) operating in coastal areas or in the archipelago. We prop...
By Bertil Grelsson, Andreas Robinson, Michael Felsberg, Fahad Shahbaz Khan
arXiv:2609.07511v1 Announce Type: cross
Abstract: Autonomous driving relies on High Definition (HD) maps for safe navigation. Traditional HD maps construction is costly in hardware, data and human re...
By Clara Gomez, Alberto Jaenal, Antonio Artu\~nedo, Jorge Godoy, Jorge Villagra
FlexMap is a vectorized high‑definition map construction framework that works with flexible camera configurations without needing calibrated rigs or explicit 2D‑to‑BEV transformations. It replaces geometric projection with a geometry foundation model that encodes cross‑view 3D structure, and uses a spatial‑temporal enhancement module and a camera‑aware decoder to separate spatial reasoning from temporal aggregation. Experiments on nuScenes and Argoverse 2 show that FlexMap outperforms pose‑dependent baselines and remains accurate even when camera views are missing or pose estimates are inaccurate.
By Run Wang, Chaoyi Zhou, Amir Salarpour, Xi Liu, Zhi-Qi Cheng, Feng Luo, Mert D. Pes\'e, Siyu Huang