arXiv Computer Vision
Aug 27

SPARK-SAM: Learning How to Prompt and Respond for Infrared Small Target Segmentation

SPARK‑SAM is a new approach that adapts the Segment‑Anything Model (SAM) for infrared small‑target segmentation by learning target‑domain response knowledge and conditioning the decoder with an image‑conditioned joint self‑prompt state. In experiments on three IRSTD benchmarks, SPARK‑SAM achieves IoU scores of 75.78%, 86.49%, and 68.34% with only 0.726 M additional parameters, outperforming 14 retrained SAM variants. The method combines benchmark‑mask supervision with reliability‑aware response guidance, and ablations show consistent accuracy gains from response guidance and high‑resolution prompt refinement.

By Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu
Hugging Face Trending Papers
Jul 2

Boosting Infrared Small Target Detection via Logit-Domain Contrast and Adaptive Shape Refinement

Infrared small target detection (IRSTD) remains challenging due to tiny target size, low signal-to-noise ratio, severe foreground-background imbalance, and blurred boundaries in complex scenes. Existing methods usually rely on post-activation probability-domain supervision for discrimination, where weak targets and strong clutter may produce saturated and close probabilities, limiting weak-target discrimination.

arXiv Computer Vision
Aug 28

DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations

The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.

By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
arXiv Computer Vision
Sep 3

RGB-to-IR image translation for infrared vehicle detection in unseen UAV domains

The paper explores using generative models to translate RGB UAV images into synthetic infrared (IR) images for training vehicle detectors in domains where real IR data is scarce. Various translators—supervised GANs, ControlNet-based diffusion models, and LoRA-ed foundation models—were trained on paired RGB-IR datasets and applied to unseen target datasets to generate synthetic IR data. The synthetic IR images, especially those produced by Stable Diffusion 3.5 with ControlNet, significantly improved detection performance on unseen IR test sets, outperforming RGB and grayscale baselines and narrowing the gap to real IR data.

By Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden, Elfi I. S. Hofmeijer, Sebastiaan P. Snel, Klamer Schutte, Friso G. Heslinga