Infrared small target detection (IRSTD) remains challenging due to tiny target size, low signal-to-noise ratio, severe foreground-background imbalance, and blurred boundaries in complex scenes. Existing methods usually rely on post-activation probability-domain supervision for discrimination, where weak targets and strong clutter may produce saturated and close probabilities, limiting weak-target discrimination.
arXiv:2607. 04603v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamental task in remote sensing.
By Tianfang Zhang, Fengyi Wu, Lei Li, Chang Liu, Zhenming Peng, Huaping Zhang, Xiangyang Ji
arXiv:2404. 10034v3 Announce Type: replace-cross Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-level labels.
By Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Eric Granger
ADGNet introduces an Asymmetric Dual-text Guided Network for infrared small target detection, addressing challenges of pixel-level methods and multimodal approaches that lack regional guidance. It employs an Asymmetric Dual-text Prompt (ADP) with an abstract target prompt and a detailed background prompt, and an Asymmetric Dual-Branch Interaction (ADBI) module to guide visual features separately, followed by an Adaptive Feature Aggregation (AFA) module for dynamic fusion. The authors also create an Asymmetric Image-Text Infrared (AITIR) dataset with asymmetric text annotations for three public datasets, and show that ADGNet outperforms 21 state‑of‑the‑art methods.
By Tongtong Wang, Mingzhu Xu, Chenglong Yu, Jing Wang, Xiaohui Lin, Weili Guan
The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.
By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
The paper introduces Aligned Consensus Teacher (ACT), a semi‑supervised framework for visible‑infrared object detection that operates under an image‑pair‑level setting with only a few labeled pairs. ACT combines Cycle‑Consistent Region Alignment, Cross‑Modal Consensus Mean‑Teacher, and Text‑Guided Cross‑Modal Instance Augmentation to address limited supervision, pseudo‑label errors, and scarce tail‑class annotations. Experiments on DroneVehicle and VEDAI demonstrate that with just 10% labeled pairs, ACT achieves 94.3% of the fully supervised mAP.
By Qi Ming, Xiaxin Yuan, Jiahuan Zhou, Jiangmeng Li, Xudong Zhao, Zhanchao Huang, Juan Fang, Shaoguang Huang, Aleksandra Pizurica