BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 03888v1 Announce Type: cross Abstract: Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effectively across tasks.
The paper presents a weakly supervised semantic segmentation approach for mapping seagrass habitats using side‑scan sonar imagery. By training a ViT‑based encoder‑decoder with image‑level labels, class activation maps are refined into pseudo‑labels and iteratively self‑trained, achieving high mean intersection‑over‑union scores (up to 89.3 %) without pixel‑level annotations. The method also benefits from self‑supervised pretraining and demonstrates generalizability in field trials.
arXiv:2608.23479v1 Announce Type: new Abstract: Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlen...
arXiv:2606. 13302v1 Announce Type: new Abstract: Wave parameters in the nearshore are crucial for coastal engineering, shoreline protection, marine hazard assessment, and coastal management for climate resilience.
SonarLLM is a multimodal large language model that treats sonar as a native perceptual modality, combining a sonar‑specific encoder, physics‑aware feature enhancement, and reliability‑aware hierarchical fusion to align acoustic structure with optical semantics. The authors introduce SonarBench, a benchmark covering recognition, counting, visual question answering, and captioning across sonar‑only, optical‑only, and fusion settings, enabling controlled measurement of cross‑modal complementarity. SonarLLM achieves 72.0% macro accuracy on sonar‑only tasks and 68.7% under fusion, outperforming baselines by significant margins and demonstrating increasing fusion gains as optical visibility degrades.
arXiv:2608. 19710v1 Announce Type: cross Abstract: Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur.