The paper evaluates out‑of‑the‑box object detection models for automatic target detection and recognition (ATD/R) in military settings. Six YOLO variants and two DETR variants were benchmarked on a new military dataset featuring vehicles, occlusions, and small targets, with performance measured in mAP@0.5 and mAP@0.5:0.95 across air‑to‑ground and ground‑to‑ground perspectives. Findings show larger models and DETR-based approaches perform best, fine‑tuning on the VisDrone dataset improves air‑to‑ground and small‑object performance, yet all models still struggle with small targets in air‑to‑ground scenarios.
By Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccol\`o Camarlinghi, H{\aa}vard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf
arXiv:2608.29783v1 Announce Type: new
Abstract: Industrial anomaly detection is a critical component of modern manufacturing. Most traditional unsupervised methods rely on modelling normal feature di...
By Weifei Chen, Honghao Zhang, Zhiyuan You, Xinyi Le
arXiv:2605. 09697v3 Announce Type: replace-cross Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a severe scarcity of positive samples.
By Radhika Amar Desai, Modigari Narendra
The paper proposes a method called Restrict, Don't Retrain that enhances zero-shot aerial segmentation by using inference-time guidance from a vision‑language model (VLM). It combines a frozen foundation model that labels every pixel with two VLM queries: one to select relevant classes and another to locate small objects missed by the base model. Experiments on four aerial datasets show consistent performance gains at each stage where the base model is competent.
By Teresa DiMeola, Charles Walter, Hong Xiao
arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.
By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy
arXiv:2606. 04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes.
By Zirui Wang, Junwei Yu, Adam Yala, David M. Chan, Joseph E. Gonzalez, Trevor Darrell
arXiv:2602. 18094v2 Announce Type: replace-cross Abstract: Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID).
By Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen
arXiv:2510. 25573v2 Announce Type: replace-cross Abstract: Machine learning approaches for image classification have led to impressive advances in that field.
By Christopher T. Franck, Anne R. Driscoll, Zoe Szajnfarber, William H. Woodall
SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.
By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer
arXiv:2606. 20077v1 Announce Type: cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals.
By Wish Suharitdamrong, Tony Alex, Muhammad Awais, Sara Atito
arXiv:2509. 04009v2 Announce Type: replace-cross Abstract: Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit unintended patterns within the data, potentially leading to correct predictions based on incorrect or unintended but statistically relevant signals.
By Solha Kang, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem, Joris Vankerschaver, Francois Rameau, Utku Ozbulak
arXiv:2510. 22665v4 Announce Type: replace-cross Abstract: Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability.
By Qiwei Ma, Xukun Lu, Wang Liu, Puhong Duan, Xudong Kang, Shutao Li