arXiv:2508.13488v2 Announce Type: replace-cross
Abstract: Loop closure detection is important for simultaneous localization and mapping (SLAM), which associates current observations with historical k...
By Jingwen Yu, Jiayi Yang, Jianhao Jiao, Anjun Hu, Zhonghang Liu, Jiankun Wang, Ping Tan, Hong Zhang
arXiv:2607. 12811v1 Announce Type: cross Abstract: Although topological mapping and navigation have been studied extensively, the specific role and downstream effect of loop closures in purely topological representations has received relatively little attention.
By Sarthak Chittawar, Vansh Garg, Aditya Vadali, Krish Pandya, Rohit Jayanti, Sourav Garg, Madhava Krishna
arXiv:2409. 11972v4 Announce Type: replace-cross Abstract: Enabling robots to autonomously discover high-level spatial concepts (e.
By Jose Andres Millan-Romera, Muhammad Shaheer, Miguel Fernandez-Cortizas, Martin R. Oswald, Holger Voos, Jose Luis Sanchez-Lopez
AquaBEV is a monocular underwater occupancy model that predicts local bird’s‑eye‑view (BEV) occupancy from a single RGB image. It uses paired 3D imaging sonar data as geometric supervision during training, mapping visual features into a calibration‑free polar representation and decoding along the range dimension before reconstructing Cartesian BEV coordinates. In a controlled underwater occupancy benchmark, AquaBEV outperforms the strongest transferred baseline with 31.4 % Visible IoU and 38.6 % Observed IoU, achieving 4.0 % and 4.3 % relative improvements respectively.
By Trung Tien Dong, Shengji Jin, Chen Chen, Yi Sheng, Xiaomin Lin
arXiv:2608.29433v1 Announce Type: new
Abstract: Sonars generate a significant amount of noise. With the advent of new technology capable of producing full 3D point clouds, the noise is amplified in s...
By Aditya Penumarti, Khanh Dong, Zi-Hao Zhang, Yongkyoon Park, Zhenqi Wu, Trung Dong, Shahriar Negahdaripour, Xiaomin Lin, Jane Shin
uScenes is a new multimodal dataset for underwater robot perception that provides synchronized 3D multibeam sonar point clouds and RGB imagery. It comprises 110 scenes with 95,834 observations, totaling 277.6 minutes of data collected during multiple field sessions. The dataset aims to support research in underwater sensor fusion, cross‑modal representation learning, and 3D scene understanding.
By Trung Tien Dong, Zhenqi Wu, Aditya Penumarti, Zi-Hao Zhang, Micaiah Bartlett, Jane Shin, Xiaomin Lin
arXiv:2510. 04100v2 Announce Type: replace-cross Abstract: Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standardized evaluation metrics, datasets, and protocols.
By Jiaming Wang, Jizhuo Chen, Diwen Liu, Harold Soh
arXiv:2609.17421v1 Announce Type: new
Abstract: Audio-Visual Navigation (AVN) requires an agent to localize and navigate toward a continuously vocalizing target relying solely on visual observations...
By Shaohang Wu, Yinfeng Yu
arXiv:2609.25271v1 Announce Type: cross
Abstract: Sidescan sonar is a common sensor for both manned and autonomous marine exploration and mapping, yet very few methods build upon or exploit the geome...
By Kalin Norman, Joshua G. Mangelson
Researchers combined an efficient algorithm with dedicated hardware to rapidly generate 3D maps for navigation using minimal memory and power.
By Adam Zewe | MIT News
arXiv:2607.11099v2 Announce Type: replace-cross
Abstract: Reliable visual data association is fundamental to visual SLAM (V-SLAM), as it directly determines the quality of the camera pose estimation...
By Ting-Wei Ou, Huang-Ting Lin, Kuu-Young Young
arXiv:2605. 06317v4 Announce Type: replace-cross Abstract: Existing Vision-Language Navigation (VLN) methods typically adopt an egocentric, step-by-step paradigm, which struggles with error accumulation and limits efficiency.
By Dijia Zhan, Jinyi Li, Chenxi Zheng, Shaoyu Huang, Yong Li, Jie Tang, Xuemiao Xu