SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.
By Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
arXiv:2608. 19710v1 Announce Type: cross Abstract: Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur.
By Mohammad Arif Ul Alam
arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.
By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
The paper introduces SASC-USOD, a framework for underwater salient object detection that learns spatially adaptive coordination between two structural representations: a boundary-sensitive representation using Laplacian filtering and a region-coherent representation via dual-range anisotropic large-kernel aggregation. A spatial coordination module estimates the relative reliability of these representations and adaptively blends them based on image content. Experiments on USOD10K and USOD benchmarks show that SASC-USOD outperforms existing methods, reducing MAE by 4.07% and 23.53% respectively, and its lightweight variant achieves 21 FPS on an NVIDIA Jetson TX2 NX.
By Lin Hong, Chenhui Wang, Linan Deng, Yuning Cui, Yu Zhang, Xin Wang, Bojian Zhang, Xingchen Yang, Fumin Zhang
arXiv:2606. 10819v1 Announce Type: cross Abstract: RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery.
By Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
arXiv:2608.20944v1 Announce Type: new
Abstract: Multimodal object detection in remote sensing faces challenges due to semantic heterogeneity and modality-specific noise interference. To this end, we...
By Xin Wu, Zhenyu Gao, Qiankun Zhang, Shaoyong Guo
arXiv:2606. 29136v1 Announce Type: cross Abstract: Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of the conventional RGB frames.
By Yu Li, Yuenan Hou, Yingmei Wei, Jiangming Chen, Yanming Guo
arXiv:2608.22072v1 Announce Type: new
Abstract: Autonomous underwater vehicles (AUVs) are increasingly important tools in industries ranging from research, to energy, to defense. AUVs are power-const...
By Gwenevere Frank, Gert Cauwenberghs
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2604.00313v3 Announce Type: replace
Abstract: Underwater image classification is constrained by the cost of annotation and by the computational and methodological requirements of task-specific...
By Thomas Manuel Rost, Martina Figlia, F. Morgado-Dias, Marko Radeta
Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks.
arXiv:2607. 00746v1 Announce Type: cross Abstract: The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception.
By Xiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang, Linna Song, Qingliang Luo, Chufan Guo, Kuifeng Su