arXiv:2607. 23770v1 Announce Type: new Abstract: In this work we study Automatic Target Recognition (ATR) for Synthetic Aperture Sonar (SAS) data with a focus on deep neural networks (DNNs).
By C. J. Moore, Gregory D. Vetaw, Jordan Malof
The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.
By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw
arXiv:2608.22072v1 Announce Type: new
Abstract: Autonomous underwater vehicles (AUVs) are increasingly important tools in industries ranging from research, to energy, to defense. AUVs are power-const...
By Gwenevere Frank, Gert Cauwenberghs
The paper examines whether standard image‑fidelity metrics (SSIM, PSNR, MSE) accurately reflect the performance of GAN‑generated synthetic sonar data in robotic perception tasks. Using a Pix2Pix GAN with four discriminator configurations (PixelGAN, PatchGAN‑16, PatchGAN‑70, ImageGAN), the authors train object detectors (YOLOX‑S, YOLOX‑L, Faster R‑CNN) solely on real sonar images and evaluate them on the synthetic outputs. Results show a mismatch: the discriminator that yields the best pixel‑level scores does not always produce the best detection performance, with PatchGAN models achieving strong downstream results despite lower SSIM/PSNR/MSE values.
By Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns
The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.
By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
The study evaluates lesion‑guided region‑of‑interest (ROI) deep learning for ovarian ultrasound classification, comparing it to global image, lesion contour, and contour‑based radiomics approaches across two public datasets. Using four deep‑learning architectures, the lesion‑guided ROI strategy achieved the highest accuracy (93.10% on MMOTU and 97.56% on OUD) with an AUC of 0.99, while requiring less annotation effort than contour‑based methods.
By Mehran Ahmad, Ali Abbasian Ardakani, Afshin Mohammadi, Alisa Mohebbi, Gernot Kronreif, Sepideh Hatamikia
arXiv:2608. 09996v1 Announce Type: cross Abstract: Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis.
By Samar Garrab, Ghada Achour
The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.
By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv:2511. 21325v2 Announce Type: replace-cross Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs.
By Ido Nitzan Hidekel, Gal lifshitz, Khen Cohen, Dan Raviv
arXiv:2407.03463v2 Announce Type: replace-cross
Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...
By Jes\'us M Rodr\'iguez-de-Vera, Imanol G Estepa, Ignacio Saras\'ua, Bhalaji Nagarajan, Petia Radeva
The paper presents a weakly supervised semantic segmentation approach for mapping seagrass habitats using side‑scan sonar imagery. By training a ViT‑based encoder‑decoder with image‑level labels, class activation maps are refined into pseudo‑labels and iteratively self‑trained, achieving high mean intersection‑over‑union scores (up to 89.3 %) without pixel‑level annotations. The method also benefits from self‑supervised pretraining and demonstrates generalizability in field trials.
By Hayat Rajani, Nuno Gracias, Rafael Garcia
arXiv:2410. 16089v2 Announce Type: replace Abstract: The unique cost, flexibility, speed, and efficiency of modern UAVs make them an attractive choice in many applications in contemporary society.
By Nikos Sakellariou (Centre for Research and Technology Hellas, Information Technologies Institute), Antonios Lalas (Centre for Research and Technology Hellas, Information Technologies Institute), Konstantinos Votis (Centre for Research and Technology Hellas, Information Technologies Institute), Dimitrios Tzovaras (Centre for Research and Technology Hellas, Information Technologies Institute)