arXiv AI

From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection

arXiv Machine Learning
Sep 4

A Real-Calibrated Synthetic-First Data Engine

The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.

By Yukang Shen, Zhiguo Liu, Yingshu Li, Yan Huang
arXiv Computer Vision
Sep 23

Semantically-Guided Domain Randomization for Industrial Object Detection in Low-Image-Budget Regimes

Semantically-Guided Domain Randomization (S‑GDR) is an annotation‑free pipeline that uses vision‑language model captioning of a small real reference set, diffusion‑based background synthesis, and mask‑based object composition to generate synthetic training data. In a high‑mix, low‑volume automotive detection benchmark, S‑GDR achieves a mAP50‑95 of 0.739 with only 200 synthetic images, outperforming a domain‑randomized render baseline and several other synthetic data methods under the same budget. These results suggest S‑GDR is a viable alternative for training visual perception systems when annotation, energy, and time resources are severely limited.

By Jose Moises Araya-Martinez, Gautham Mohan, Jens Lambrecht
arXiv Computer Vision
Sep 3

Domain shift-robust object detection with GenAI image editing

The paper investigates using diffusion-based generative image editing to improve object detector robustness against domain shifts, specifically camouflaged military vehicle detection. By synthetically adding foliage, netting, and multi‑spectral camouflage to training data with models such as Qwen Image Edit 2509 and Flux.2 Dev, the authors demonstrate significant mAP gains (up to +20.1 for foliage) over detectors trained on uncamouflaged data. LoRA fine‑tuning further boosts performance for the more challenging multi‑spectral camouflage.

By Isabel D. Stein, Thijs A. Eker, Sebastiaan P. Snel, Ella P. Fokkinga, Klamer Schutte, Luca Ambrogioni, Friso G. Heslinga
arXiv Machine Learning
Aug 28

A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation

The paper introduces a framework that adapts a diffusion model to a target urban domain using only imperfect pseudo‑labels, enabling the generation of high‑fidelity, target‑aligned images from semantic maps of any synthetic dataset. By filtering poor generations, correcting image‑label misalignments, and standardising semantics, the method transforms low‑effort synthetic data into competitive real‑domain training sets. Experiments on five synthetic and two real datasets show up to +8.0 %pt mIoU improvement over state‑of‑the‑art translation methods, demonstrating that rapidly constructed synthetic datasets can match the performance of high‑effort, manually designed ones.

By Damjan Kal\v{s}an, Denis Zavadski, Tim K\"uchler, Haebom Lee, Stefan Roth, Carsten Rother