SPARK‑SAM is a new approach that adapts the Segment‑Anything Model (SAM) for infrared small‑target segmentation by learning target‑domain response knowledge and conditioning the decoder with an image‑conditioned joint self‑prompt state. In experiments on three IRSTD benchmarks, SPARK‑SAM achieves IoU scores of 75.78%, 86.49%, and 68.34% with only 0.726 M additional parameters, outperforming 14 retrained SAM variants. The method combines benchmark‑mask supervision with reliability‑aware response guidance, and ablations show consistent accuracy gains from response guidance and high‑resolution prompt refinement.
By Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu
arXiv:2608. 05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains.
By Aohua Li, Jin Kuang, Yubing Lu, Pingping Liu
arXiv:2609.38111v1 Announce Type: new
Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when...
By Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng
Infrared small target detection (IRSTD) remains challenging due to tiny target size, low signal-to-noise ratio, severe foreground-background imbalance, and blurred boundaries in complex scenes. Existing methods usually rely on post-activation probability-domain supervision for discrimination, where weak targets and strong clutter may produce saturated and close probabilities, limiting weak-target discrimination.
The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.
By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
The paper explores using generative models to translate RGB UAV images into synthetic infrared (IR) images for training vehicle detectors in domains where real IR data is scarce. Various translators—supervised GANs, ControlNet-based diffusion models, and LoRA-ed foundation models—were trained on paired RGB-IR datasets and applied to unseen target datasets to generate synthetic IR data. The synthetic IR images, especially those produced by Stable Diffusion 3.5 with ControlNet, significantly improved detection performance on unseen IR test sets, outperforming RGB and grayscale baselines and narrowing the gap to real IR data.
By Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden, Elfi I. S. Hofmeijer, Sebastiaan P. Snel, Klamer Schutte, Friso G. Heslinga
The paper explores soft prompting for few‑shot object detection with vision‑language models, showing that optimizing a small number of continuous prompt tokens—especially when placed at the cross‑modal boundary and initialized from an empty space token—can match LoRA performance while training far fewer parameters. Soft prompting also avoids catastrophic forgetting, transfers to newer models, and can be verbalized into readable prompts. The study extends these findings to manipulation tasks, indicating that VLMs already contain much of the necessary knowledge for specialized domains, and the main challenge is learning how to ask for it.
By Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan
The paper investigates why vision‑language models like LLaVA‑1.5‑7B hallucinate objects in captions and proposes a targeted fix. By ranking attention heads whose image attention drops around hallucinated words, the authors identify 32 key heads and apply a head‑sliced LoRA adapter plus an inference‑time grounding controller. On COCO images, this combined method reduces hallucinated captions from 37% to 23% and hallucinated object mentions from 15.6% to 9.6%, while also lowering object recall.
By Armaan Sandhu, Abhilasha Senapati, Hima Kammachi
arXiv:2505.03380v2 Announce Type: replace
Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains...
By Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li
The paper introduces SAM3-LoRA, a parameter‑efficient adaptation of the SAM3 segmentation foundation model using Low‑Rank Adaptation for multi‑class structural defect segmentation. It presents a supervision method that trains the model directly from COCO‑style instance segmentation by using category names as prompts, eliminating the need for prompt templates or learned embeddings. The authors also identify a failure mode where the model responds to any prompt due to positive‑prompt‑only annotations and resolve it with exhaustive hard‑negative prompting, achieving significant improvements in pixel IoU and instance recall on both a tunnel lining dataset and the Structural Defects Dataset.
By P. Malaisree, S. Youwai, S. Janrungautai, D. Amorndechaphon, P. Rojanavasu, W. Songkitti
arXiv:2607. 24354v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results.
By Haoyue Liu, Xiaoyu Ma, Ye Chen, Yuexian Zou, Xiaoying Tang
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can be unreliable in VIS-IR data because imaging-physics differences make some spatially paired regions inherently less comparable, and aligning them with equal strength hinders representation learning and downstream transfer.