RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.
arXiv:2608. 19710v1 Announce Type: cross Abstract: Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur.
arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.
The paper introduces SASC-USOD, a framework for underwater salient object detection that learns spatially adaptive coordination between two structural representations: a boundary-sensitive representation using Laplacian filtering and a region-coherent representation via dual-range anisotropic large-kernel aggregation. A spatial coordination module estimates the relative reliability of these representations and adaptively blends them based on image content. Experiments on USOD10K and USOD benchmarks show that SASC-USOD outperforms existing methods, reducing MAE by 4.07% and 23.53% respectively, and its lightweight variant achieves 21 FPS on an NVIDIA Jetson TX2 NX.
arXiv:2606. 10819v1 Announce Type: cross Abstract: RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery.
arXiv:2608.20944v1 Announce Type: new Abstract: Multimodal object detection in remote sensing faces challenges due to semantic heterogeneity and modality-specific noise interference. To this end, we...