arXiv:2608.20754v1 Announce Type: new
Abstract: Promptable segmentation models provide a reusable interface, but direct transfer to automatic infrared small-target segmentation (IRSTD) exposes a mism...
By Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu
arXiv:2608. 05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains.
By Aohua Li, Jin Kuang, Yubing Lu, Pingping Liu
arXiv:2609.38111v1 Announce Type: new
Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when...
By Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng
Infrared small target detection (IRSTD) remains challenging due to tiny target size, low signal-to-noise ratio, severe foreground-background imbalance, and blurred boundaries in complex scenes. Existing methods usually rely on post-activation probability-domain supervision for discrimination, where weak targets and strong clutter may produce saturated and close probabilities, limiting weak-target discrimination.
The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.
By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
The paper explores soft prompting for few‑shot object detection with vision‑language models, showing that optimizing a small number of continuous prompt tokens—especially when placed at the cross‑modal boundary and initialized from an empty space token—can match LoRA performance while training far fewer parameters. Soft prompting also avoids catastrophic forgetting, transfers to newer models, and can be verbalized into readable prompts. The study extends these findings to manipulation tasks, indicating that VLMs already contain much of the necessary knowledge for specialized domains, and the main challenge is learning how to ask for it.
By Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez, Wei Zhao, Wolfgang M. Pauli, John Galeotti, Deva Ramanan