arXiv Computer Vision

Hull First, Wake Second: Wake-Reliance Suppression for Robust Maritime Vessel Detection

HullWake is a new maritime vessel detection framework that prioritizes hull detection before wake analysis to address the wake-reliance problem. It separates hull evidence from wake context, extracts wake cues via bidirectional proposal-anchored corridors, and suppresses wake-dominant predictions through multiple supervisory strategies. The method is evaluated on a wake-oriented dataset and demonstrates improvements in overall AP, robustness to weak or no-wake vessels, reduction of wake-like false positives, and stability of confidence after wake attenuation.

arXiv AI
Sep 21

LoRA Enhanced Contrastive Learning with SAS Vision Transformers

The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.

By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw
arXiv Computer Vision
Sep 24

Strip Convolution and Direction-Aware Exclusion Loss for Oriented Ship Detection

The paper introduces a new oriented ship detector that combines a C3k2_Strip module, which uses orthogonal strip convolutions to better capture elongated hull structures, with a Class-Aware Direction-Aware Exclusion Loss (CA-DAEL) that suppresses redundant predictions by leveraging class, direction, and confidence cues. Experiments on HRSC2016 and DIOR-R datasets show the method achieving 78.45% and 53.71% mAP50:95, respectively, with only 2.91M parameters. On HRSC2016, the approach outperforms the YOLOv11-OBB baseline by 6.32 percentage points in mAP50:95, highlighting its effectiveness for accurate oriented ship detection.

By Bin Chen, Yuanyuan Liu, Peng Yang, Chao Lu
arXiv Computer Vision
Aug 25

Learning Spatially Adaptive Structural Coordination for Underwater Salient Object Detection

The paper introduces SASC-USOD, a framework for underwater salient object detection that learns spatially adaptive coordination between two structural representations: a boundary-sensitive representation using Laplacian filtering and a region-coherent representation via dual-range anisotropic large-kernel aggregation. A spatial coordination module estimates the relative reliability of these representations and adaptively blends them based on image content. Experiments on USOD10K and USOD benchmarks show that SASC-USOD outperforms existing methods, reducing MAE by 4.07% and 23.53% respectively, and its lightweight variant achieves 21 FPS on an NVIDIA Jetson TX2 NX.

By Lin Hong, Chenhui Wang, Linan Deng, Yuning Cui, Yu Zhang, Xin Wang, Bojian Zhang, Xingchen Yang, Fumin Zhang
arXiv Computer Vision
Sep 23

Observer Choice and Threshold Selection in Retinal Vessel Segmentation: A Subject-Separated Evaluation

The study investigates how the choice of annotation used to set a segmentation threshold influences retinal vessel segmentation performance. Using all 28 CHASE DB1 images and two human observers, the authors fit random forests and Extra Trees models, then compare five threshold policies—including fixed, observer‑tuned, mean‑observer, and maximin tuning—on the same score maps. Results show that maximin tuning alters thresholds in most fits but yields negligible changes in worst‑observer Dice scores, suggesting no accuracy advantage in this cohort.

By Wenhao Xu, Yixian Kong, Ting Pan, Changwei Wang, Feilong Wang, Rongtao Xu
arXiv Computer Vision
Sep 3

KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection

KSG‑Net introduces a Key‑Sparse and Global‑Context learning framework for maritime 3D ship detection, addressing weak feature representation of small, sparse vessels and limited global modeling of large vessels. The network employs a Key Sparse Multi‑scale Aggregation module to select informative voxels and aggregate cross‑scale features, and a Global Context Aggregation module to capture long‑range geometric dependencies via scene‑level context modeling. Experiments on the Thames River vessel dataset and simulated data show that KSG‑Net outperforms existing methods in multi‑scale vessel detection and remains robust in complex maritime environments.

By Zhouyuan Huai, Meiqi Wan, Yan Yang, Minshi Chen, Xin Yuan, Wei Wang, Xiao Wang