arXiv Computer Vision

DGNet: Dual-knowledge Guided Network for Infrared Small Target Detection

arXiv AI
Sep 2

ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection

ADGNet introduces an Asymmetric Dual-text Guided Network for infrared small target detection, addressing challenges of pixel-level methods and multimodal approaches that lack regional guidance. It employs an Asymmetric Dual-text Prompt (ADP) with an abstract target prompt and a detailed background prompt, and an Asymmetric Dual-Branch Interaction (ADBI) module to guide visual features separately, followed by an Adaptive Feature Aggregation (AFA) module for dynamic fusion. The authors also create an Asymmetric Image-Text Infrared (AITIR) dataset with asymmetric text annotations for three public datasets, and show that ADGNet outperforms 21 state‑of‑the‑art methods.

By Tongtong Wang, Mingzhu Xu, Chenglong Yu, Jing Wang, Xiaohui Lin, Weili Guan
arXiv Computer Vision
Aug 28

DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations

The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.

By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
Hugging Face Trending Papers
Jul 26

ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and downstream tasks. However, existing methods typically rely on predefined discrete control conditions, leading to a sparse space that fails to support fine-grained modulation demands.

Hugging Face Trending Papers
Jul 2

Boosting Infrared Small Target Detection via Logit-Domain Contrast and Adaptive Shape Refinement

Infrared small target detection (IRSTD) remains challenging due to tiny target size, low signal-to-noise ratio, severe foreground-background imbalance, and blurred boundaries in complex scenes. Existing methods usually rely on post-activation probability-domain supervision for discrimination, where weak targets and strong clutter may produce saturated and close probabilities, limiting weak-target discrimination.

Hugging Face Trending Papers
Aug 13

P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.

arXiv Computer Vision
Sep 23

C2FXNet: Coarse-to-Fine Scene Expert for Unified Object Detection across Adverse Weather

C2FXNet is a unified object detection framework designed for adverse weather conditions. It uses a dual-level guidance mechanism: a Multi-step Reasoning Router (MRR) for coarse scene reasoning and a Fine Scene Refinement (FSR) module for fine-grained semantic adjustment. A Scene-aware Mixture-of-Experts (SMoE) dynamically combines scene-specific experts, enabling robust detection across foggy, dark, and clear scenes without scene-specific training.

By Tianle Fang, Zhenbing Liu, Chong Yin, Bolun Li, Haoxiang Lu