arXiv Computer Vision

Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation

arXiv Computer Vision
Sep 16

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.

By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
arXiv Machine Learning
Jun 2

Domain Adaptation with a Single Vision-Language Embedding

arXiv:2410. 21361v2 Announce Type: replace-cross Abstract: Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions.

By Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick P\'erez, Raoul de Charette
arXiv Computer Vision
Sep 24

ICM: Intra-class Mixing for Domain Adaptation in Adverse Weather

The paper introduces ICM, an Intra-Class Mixing Consistency framework for unsupervised domain adaptation in semantic segmentation under adverse weather. ICM mixes regions within the same image and semantic class to maintain realistic layouts, contrasting prior methods that combine across images or domains. On the Cityscapes → ACDC benchmark, ICM achieves 75.7% mIoU, surpassing previous state‑of‑the‑art results by 1.9 percentage points.

By Boying Li, Chang Liu, Britta Ayano Wilde, Gy\"orgy Kov\'acs, Tosin Adewumi, Bj\"orn Backe, Hamam Mokayed
arXiv Machine Learning
Aug 4

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

arXiv:2604. 15622v3 Announce Type: replace-cross Abstract: Always-on contextual AI runs language-aligned vision foundation models (VFMs) on edge devices, where the on-device model is the dominant continuous compute cost under strict latency and power limits.

By Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li
arXiv Computer Vision
Aug 27

Difficulty-Aware Sample Allocation for Adaptive Data Augmentation in Semantic Segmentation

The paper proposes Difficulty-Aware Sample Allocation (DASA), a framework that assigns stronger data augmentation to training samples deemed more difficult based on a composite difficulty score. This score integrates prediction ambiguity, training loss, class rarity, and boundary complexity, and is used to modulate augmentation strength during iterative training. Experiments on Oxford‑IIIT Pet and binary Pascal VOC with U‑Net, DeepLabV3, and SegFormer‑B0 demonstrate that DASA outperforms standard training and matches or exceeds single‑signal adaptive baselines, notably raising DeepLabV3’s mIoU from 0.633 to 0.740 on Oxford‑IIIT Pet.

By Olasimbo Ayodeji Arigbabu, Abimbola Ismail Arigbabu