Concentration, Not Uncertainty: Why Targeted Synthetic Data Doesn't Help Camouflaged Object Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.38010v1 Announce Type: cross Abstract: Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety m...
arXiv:2607. 02637v1 Announce Type: cross Abstract: Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models.
Generative Verification introduces an active learning strategy for object detection that uses an independent generative model to re‑derive a detection’s label from the pixels inside its predicted box. The disagreement between the detector’s label and the verifier’s label serves as the acquisition signal, automatically combining localization and classification errors into a single scalar and eliminating the need for hand‑weighted terms. Experiments on PASCAL VOC and MS‑COCO show that this signal outperforms traditional uncertainty and ensemble criteria, improving mAP50 by about one point per round, especially in early rounds where confident detector errors are most common.
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
The paper demonstrates that a targeted adversarial perturbation can reduce a vision‑language model’s training loss to near zero for a fixed target caption, yet the same model, when generating freely, still produces the correct description. This phenomenon, termed the train/inference gap, is traced to a single autoregressive step where the target token’s rank is fixed across all images, and further analysis shows that the language decoder, rather than the visual encoder, determines whether the corrupted signal is amplified or suppressed. The study uses a controlled two‑stage PGD attack on Qwen2.5‑VL‑7B‑Instruct and evaluates the effect on 200 held‑out COCO images, revealing that adversarial robustness in autoregressive VLMs largely depends on the language decoder’s prior. whyItMatters":"The findings suggest that defenses and faithfulness evaluations for deployed vision‑language models should focus on the language decoder rather than the visual encoder, as the former is the key determinant of robustness to adversarial perturbations."
The paper investigates using diffusion-based generative image editing to improve object detector robustness against domain shifts, specifically camouflaged military vehicle detection. By synthetically adding foliage, netting, and multi‑spectral camouflage to training data with models such as Qwen Image Edit 2509 and Flux.2 Dev, the authors demonstrate significant mAP gains (up to +20.1 for foliage) over detectors trained on uncamouflaged data. LoRA fine‑tuning further boosts performance for the more challenging multi‑spectral camouflage.