What is the Added Value of UDA in the VFM Era?
arXiv:2504.18190v2 Announce Type: replace Abstract: Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source d...
The paper introduces a cross‑modal pseudo‑labeling pipeline for unsupervised domain adaptation in semantic segmentation, particularly for waste sorting. It combines SAM for class‑agnostic region proposals with EVA‑CLIP to assign semantic labels via region‑text similarity, applying confidence filtering to ensure reliable pseudo‑labels for self‑training. An optional BLIP‑based language‑grounded verification further refines ambiguous regions, and the method shows consistent improvements over source‑only baselines on synthetic‑to‑real driving and lab‑to‑factory waste sorting shifts.
arXiv:2504.18190v2 Announce Type: replace Abstract: Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source d...
Semantically-Guided Domain Randomization (S‑GDR) is an annotation‑free pipeline that uses vision‑language model captioning of a small real reference set, diffusion‑based background synthesis, and mask‑based object composition to generate synthetic training data. In a high‑mix, low‑volume automotive detection benchmark, S‑GDR achieves a mAP50‑95 of 0.739 with only 200 synthetic images, outperforming a domain‑randomized render baseline and several other synthetic data methods under the same budget. These results suggest S‑GDR is a viable alternative for training visual perception systems when annotation, energy, and time resources are severely limited.
arXiv:2410. 21361v2 Announce Type: replace-cross Abstract: Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions.
The paper introduces a framework that adapts a diffusion model to a target urban domain using only imperfect pseudo‑labels, enabling the generation of high‑fidelity, target‑aligned images from semantic maps of any synthetic dataset. By filtering poor generations, correcting image‑label misalignments, and standardising semantics, the method transforms low‑effort synthetic data into competitive real‑domain training sets. Experiments on five synthetic and two real datasets show up to +8.0 %pt mIoU improvement over state‑of‑the‑art translation methods, demonstrating that rapidly constructed synthetic datasets can match the performance of high‑effort, manually designed ones.
Retraining visual perception pipelines in High-Mix, Low-Volume (HMLV) automotive manufacturing must be carried out under tight annotation, energy, and time budgets, yet most Synthetic Data Generation...
The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe perf...
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, enabling attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate this method, especially for sparse, real‑world differences, and the authors provide open‑weight models and code for reproducibility.
The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, allowing attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate these methods, especially for sparse, real‑world differences, with open‑weight models to ensure reproducibility.
The paper introduces Self‑Evolutionary CLIP (SE‑CLIP), a semi‑supervised framework that adapts vision‑language models like CLIP to satellite imagery. SE‑CLIP uses a two‑phase pipeline: an initial warm‑up on a small set of annotated seeds followed by a recursive discovery phase that iteratively selects high‑confidence samples from unlabeled data. A class‑balanced selection strategy is applied to keep the evolving support set balanced, and experiments on the UCM and NWPU benchmarks show that SE‑CLIP outperforms existing semi‑supervised methods.
arXiv:2504.14280v2 Announce Type: replace-cross Abstract: As machine learning evolves, domain generalization (DG) and domain adaptation (DA) have become crucial for improving model robustness across...