arXiv Computer Vision

Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

The paper investigates Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) imagery, comparing modern convolutional neural networks (CNNs) and transformer-based deep neural networks (DNNs). It examines how factors such as network size, architecture, pretraining methods, data augmentation, and regularization influence performance, aiming to identify the highest-performing model and provide a training roadmap for state‑of‑the‑art SAS‑ATR systems.

arXiv AI
Sep 21

LoRA Enhanced Contrastive Learning with SAS Vision Transformers

The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.

By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw
arXiv Machine Learning
Sep 17

Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception

The paper examines whether standard image‑fidelity metrics (SSIM, PSNR, MSE) accurately reflect the performance of GAN‑generated synthetic sonar data in robotic perception tasks. Using a Pix2Pix GAN with four discriminator configurations (PixelGAN, PatchGAN‑16, PatchGAN‑70, ImageGAN), the authors train object detectors (YOLOX‑S, YOLOX‑L, Faster R‑CNN) solely on real sonar images and evaluate them on the synthetic outputs. Results show a mismatch: the discriminator that yields the best pixel‑level scores does not always produce the best detection performance, with PatchGAN models achieving strong downstream results despite lower SSIM/PSNR/MSE values.

By Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns
arXiv Machine Learning
Sep 25

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv Computer Vision
Aug 27

Less Contouring, More Accuracy: Lesion-Guided ROI Deep Learning for Ovarian Ultrasound Classification

The study evaluates lesion‑guided region‑of‑interest (ROI) deep learning for ovarian ultrasound classification, comparing it to global image, lesion contour, and contour‑based radiomics approaches across two public datasets. Using four deep‑learning architectures, the lesion‑guided ROI strategy achieved the highest accuracy (93.10% on MMOTU and 97.56% on OUD) with an AUC of 0.99, while requiring less annotation effort than contour‑based methods.

By Mehran Ahmad, Ali Abbasian Ardakani, Afshin Mohammadi, Alisa Mohebbi, Gernot Kronreif, Sepideh Hatamikia
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv Machine Learning
Aug 26

Weakly Supervised Seafloor Segmentation for Seagrass Habitat Mapping in Side-Scan Sonar Imagery

The paper presents a weakly supervised semantic segmentation approach for mapping seagrass habitats using side‑scan sonar imagery. By training a ViT‑based encoder‑decoder with image‑level labels, class activation maps are refined into pseudo‑labels and iteratively self‑trained, achieving high mean intersection‑over‑union scores (up to 89.3 %) without pixel‑level annotations. The method also benefits from self‑supervised pretraining and demonstrates generalizability in field trials.

By Hayat Rajani, Nuno Gracias, Rafael Garcia
arXiv AI
Jun 16

Multi-Sensor Fusion for UAV Classification Based on Feature Maps of Image and Radar Data

arXiv:2410. 16089v2 Announce Type: replace Abstract: The unique cost, flexibility, speed, and efficiency of modern UAVs make them an attractive choice in many applications in contemporary society.

By Nikos Sakellariou (Centre for Research and Technology Hellas, Information Technologies Institute), Antonios Lalas (Centre for Research and Technology Hellas, Information Technologies Institute), Konstantinos Votis (Centre for Research and Technology Hellas, Information Technologies Institute), Dimitrios Tzovaras (Centre for Research and Technology Hellas, Information Technologies Institute)