arXiv Computer Vision

Concentration, Not Uncertainty: Why Targeted Synthetic Data Doesn't Help Camouflaged Object Detection

arXiv Computer Vision
Sep 18

Generative Verification: Rethinking the Uncertainty Signal for Active Learning of Object Detection

Generative Verification introduces an active learning strategy for object detection that uses an independent generative model to re‑derive a detection’s label from the pixels inside its predicted box. The disagreement between the detector’s label and the verifier’s label serves as the acquisition signal, automatically combining localization and classification errors into a single scalar and eliminating the need for hand‑weighted terms. Experiments on PASCAL VOC and MS‑COCO show that this signal outperforms traditional uncertainty and ensemble criteria, improving mAP50 by about one point per round, especially in early rounds where confident detector errors are most common.

By Licheng Zhang, Zheng Gong
arXiv AI
4d ago

Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

The paper demonstrates that a targeted adversarial perturbation can reduce a vision‑language model’s training loss to near zero for a fixed target caption, yet the same model, when generating freely, still produces the correct description. This phenomenon, termed the train/inference gap, is traced to a single autoregressive step where the target token’s rank is fixed across all images, and further analysis shows that the language decoder, rather than the visual encoder, determines whether the corrupted signal is amplified or suppressed. The study uses a controlled two‑stage PGD attack on Qwen2.5‑VL‑7B‑Instruct and evaluates the effect on 200 held‑out COCO images, revealing that adversarial robustness in autoregressive VLMs largely depends on the language decoder’s prior. whyItMatters":"The findings suggest that defenses and faithfulness evaluations for deployed vision‑language models should focus on the language decoder rather than the visual encoder, as the former is the key determinant of robustness to adversarial perturbations."

By Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama
arXiv Computer Vision
Sep 3

Domain shift-robust object detection with GenAI image editing

The paper investigates using diffusion-based generative image editing to improve object detector robustness against domain shifts, specifically camouflaged military vehicle detection. By synthetically adding foliage, netting, and multi‑spectral camouflage to training data with models such as Qwen Image Edit 2509 and Flux.2 Dev, the authors demonstrate significant mAP gains (up to +20.1 for foliage) over detectors trained on uncamouflaged data. LoRA fine‑tuning further boosts performance for the more challenging multi‑spectral camouflage.

By Isabel D. Stein, Thijs A. Eker, Sebastiaan P. Snel, Ella P. Fokkinga, Klamer Schutte, Luca Ambrogioni, Friso G. Heslinga
arXiv Machine Learning
Aug 11

From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition

arXiv:2308. 04553v4 Announce Type: replace-cross Abstract: Visual recognition models are prone to learning spurious correlations induced by a biased training set where certain conditions $B$ (\eg, Indoors) are over-represented in certain classes $Y$ (\eg, Big Dogs).

By Maan Qraitem, Kate Saenko, Bryan A. Plummer
arXiv Machine Learning
Sep 4

A Real-Calibrated Synthetic-First Data Engine

The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.

By Yukang Shen, Zhiguo Liu, Yingshu Li, Yan Huang