arXiv:2507. 18632v2 Announce Type: replace-cross Abstract: Zero-shot domain adaptation is a method for adapting a model to a target domain without utilizing target domain image data.
By Ye-Chan Kim, SeungJu Cha, Si-Woo Kim, Taewhan Kim, Dong-Jin Kim
arXiv:2509.22650v3 Announce Type: replace
Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
By Anna Kukleva, Enis Simsar, Alessio Tonioni, Muhammad Ferjad Naeem, Federico Tombari, Jan Eric Lenssen, Bernt Schiele
Semantically-Guided Domain Randomization (S‑GDR) is an annotation‑free pipeline that uses vision‑language model captioning of a small real reference set, diffusion‑based background synthesis, and mask‑based object composition to generate synthetic training data. In a high‑mix, low‑volume automotive detection benchmark, S‑GDR achieves a mAP50‑95 of 0.739 with only 200 synthetic images, outperforming a domain‑randomized render baseline and several other synthetic data methods under the same budget. These results suggest S‑GDR is a viable alternative for training visual perception systems when annotation, energy, and time resources are severely limited.
By Jose Moises Araya-Martinez, Gautham Mohan, Jens Lambrecht
arXiv:2606. 09362v1 Announce Type: cross Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues.
By Eduardo Borges, Manuel Abreu, Lu\'is Garrote, Urbano J. Nunes
arXiv:2504.18190v2 Announce Type: replace
Abstract: Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source d...
By Brun\'o B. Englert, Tommie Kerssies, Gijs Dubbelman
arXiv:2607. 02269v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG).
By Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma