Semantically-Guided Domain Randomization for Industrial Object Detection in Low-Image-Budget Regimes
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
Semantically-Guided Domain Randomization (S‑GDR) is an annotation‑free pipeline that uses vision‑language model captioning of a small real reference set, diffusion‑based background synthesis, and mask‑based object composition to generate synthetic training data. In a high‑mix, low‑volume automotive detection benchmark, S‑GDR achieves a mAP50‑95 of 0.739 with only 200 synthetic images, outperforming a domain‑randomized render baseline and several other synthetic data methods under the same budget. These results suggest S‑GDR is a viable alternative for training visual perception systems when annotation, energy, and time resources are severely limited.
arXiv:2609.38010v1 Announce Type: cross Abstract: Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety m...
arXiv:2410. 21361v2 Announce Type: replace-cross Abstract: Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions.
The paper introduces a framework that adapts a diffusion model to a target urban domain using only imperfect pseudo‑labels, enabling the generation of high‑fidelity, target‑aligned images from semantic maps of any synthetic dataset. By filtering poor generations, correcting image‑label misalignments, and standardising semantics, the method transforms low‑effort synthetic data into competitive real‑domain training sets. Experiments on five synthetic and two real datasets show up to +8.0 %pt mIoU improvement over state‑of‑the‑art translation methods, demonstrating that rapidly constructed synthetic datasets can match the performance of high‑effort, manually designed ones.
The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, enabling attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate this method, especially for sparse, real‑world differences, and the authors provide open‑weight models and code for reproducibility.
The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, allowing attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate these methods, especially for sparse, real‑world differences, with open‑weight models to ensure reproducibility.