Infrared small target detection (IRSTD) remains challenging due to tiny target size, low signal-to-noise ratio, severe foreground-background imbalance, and blurred boundaries in complex scenes. Existing methods usually rely on post-activation probability-domain supervision for discrimination, where weak targets and strong clutter may produce saturated and close probabilities, limiting weak-target discrimination.
arXiv:2607. 04603v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamental task in remote sensing.
By Tianfang Zhang, Fengyi Wu, Lei Li, Chang Liu, Zhenming Peng, Huaping Zhang, Xiangyang Ji
arXiv:2404. 10034v3 Announce Type: replace-cross Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-level labels.
By Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Eric Granger
ADGNet introduces an Asymmetric Dual-text Guided Network for infrared small target detection, addressing challenges of pixel-level methods and multimodal approaches that lack regional guidance. It employs an Asymmetric Dual-text Prompt (ADP) with an abstract target prompt and a detailed background prompt, and an Asymmetric Dual-Branch Interaction (ADBI) module to guide visual features separately, followed by an Adaptive Feature Aggregation (AFA) module for dynamic fusion. The authors also create an Asymmetric Image-Text Infrared (AITIR) dataset with asymmetric text annotations for three public datasets, and show that ADGNet outperforms 21 state‑of‑the‑art methods.
By Tongtong Wang, Mingzhu Xu, Chenglong Yu, Jing Wang, Xiaohui Lin, Weili Guan
The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.
By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
The paper introduces Aligned Consensus Teacher (ACT), a semi‑supervised framework for visible‑infrared object detection that operates under an image‑pair‑level setting with only a few labeled pairs. ACT combines Cycle‑Consistent Region Alignment, Cross‑Modal Consensus Mean‑Teacher, and Text‑Guided Cross‑Modal Instance Augmentation to address limited supervision, pseudo‑label errors, and scarce tail‑class annotations. Experiments on DroneVehicle and VEDAI demonstrate that with just 10% labeled pairs, ACT achieves 94.3% of the fully supervised mAP.
By Qi Ming, Xiaxin Yuan, Jiahuan Zhou, Jiangmeng Li, Xudong Zhao, Zhanchao Huang, Juan Fang, Shaoguang Huang, Aleksandra Pizurica
Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and downstream tasks. However, existing methods typically rely on predefined discrete control conditions, leading to a sparse space that fails to support fine-grained modulation demands.
SPARK‑SAM is a new approach that adapts the Segment‑Anything Model (SAM) for infrared small‑target segmentation by learning target‑domain response knowledge and conditioning the decoder with an image‑conditioned joint self‑prompt state. In experiments on three IRSTD benchmarks, SPARK‑SAM achieves IoU scores of 75.78%, 86.49%, and 68.34% with only 0.726 M additional parameters, outperforming 14 retrained SAM variants. The method combines benchmark‑mask supervision with reliability‑aware response guidance, and ablations show consistent accuracy gains from response guidance and high‑resolution prompt refinement.
By Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu
arXiv:2608.20870v1 Announce Type: new
Abstract: Infrared small target detection is still challenging in remote sensing imagery, because the targets are extremely small, exhibit weak local contrast, a...
By Rui Liu, Jing Nie, Ying Fu
arXiv:2606. 26973v1 Announce Type: cross Abstract: Open-set semi-supervised learning aims to leverage unlabeled data that may contain out-of-distribution outliers while maintaining performance on in-distribution classes.
By Jiahe Chen, Qian Shao, Qiyuan Chen, Jiaying He, Jintai Chen, Jian Wu, Hongxia Xu
arXiv:2608.22368v1 Announce Type: new
Abstract: While linear attention is a compelling mechanism for high-resolution object detection due to its reduced cost for global token mixing, converting the S...
By Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin, Muli Yang, Hongyuan Zhu
arXiv:2608.20754v1 Announce Type: new
Abstract: Promptable segmentation models provide a reusable interface, but direct transfer to automatic infrared small-target segmentation (IRSTD) exposes a mism...
By Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu