arXiv Computer Vision

Distance to Class Prototypes: Active Learning for Object Detection

The paper introduces a new active learning signal for object detection that relies on a supervised contrastive term added to the training objective. This term shapes an embedding space where distance reflects class membership, allowing an unlabeled detection to be scored by its distance from the predicted category’s region weighted by confidence—all from a single forward pass of one network. Experiments on PASCAL VOC and MS‑COCO show that this criterion outperforms the standard posterior and remains competitive with ensemble‑based methods while incurring only a modest 8.3% increase in parameters.

arXiv Computer Vision
Sep 18

Generative Verification: Rethinking the Uncertainty Signal for Active Learning of Object Detection

Generative Verification introduces an active learning strategy for object detection that uses an independent generative model to re‑derive a detection’s label from the pixels inside its predicted box. The disagreement between the detector’s label and the verifier’s label serves as the acquisition signal, automatically combining localization and classification errors into a single scalar and eliminating the need for hand‑weighted terms. Experiments on PASCAL VOC and MS‑COCO show that this signal outperforms traditional uncertainty and ensemble criteria, improving mAP50 by about one point per round, especially in early rounds where confident detector errors are most common.

By Licheng Zhang, Zheng Gong
arXiv AI
1d ago

Geometry-Aware Adaptation for Pretrained Models

arXiv:2307.12226v3 Announce Type: replace-cross Abstract: Machine learning models -- including prominent zero-shot models -- are often trained on datasets whose labels are only a small proportion of...

By Nicholas Roberts, Xintong Li, Dyah Adila, Sonia Cromp, Tzu-Heng Huang, Jitian Zhao, Frederic Sala
arXiv AI
Aug 19

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

The paper evaluates out‑of‑the‑box object detection models for automatic target detection and recognition (ATD/R) in military settings. Six YOLO variants and two DETR variants were benchmarked on a new military dataset featuring vehicles, occlusions, and small targets, with performance measured in mAP@0.5 and mAP@0.5:0.95 across air‑to‑ground and ground‑to‑ground perspectives. Findings show larger models and DETR-based approaches perform best, fine‑tuning on the VisDrone dataset improves air‑to‑ground and small‑object performance, yet all models still struggle with small targets in air‑to‑ground scenarios.

By Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccol\`o Camarlinghi, H{\aa}vard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf
arXiv AI
Sep 1

Background-Free Objectness Learning for Class-Agnostic Detection

Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.

By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella