arXiv:2410. 21361v2 Announce Type: replace-cross Abstract: Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions.
By Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick P\'erez, Raoul de Charette
arXiv:2607. 07179v1 Announce Type: cross Abstract: Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents.
By Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia
arXiv:2504.14280v2 Announce Type: replace-cross
Abstract: As machine learning evolves, domain generalization (DG) and domain adaptation (DA) have become crucial for improving model robustness across...
By Jindong Li, Yongguang Li, Yali Fu, Jiahong Liu, Yixin Liu, Menglin Yang, Irwin King
The paper introduces SubTTA, a test-time adaptation method for vision‑language models that aligns the semantic subspaces of visual and textual modalities to improve zero‑shot predictions. It addresses two issues: the modality gap caused by distribution shifts and visual nuisance that masks task‑specific semantics. By minimizing chordal distance between principal subspaces and projecting visual features onto a task‑specific textual subspace, SubTTA refines decision boundaries and achieves an average 2.24% improvement over existing TTA methods.
By Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong
arXiv:2511.05782v3 Announce Type: replace
Abstract: Unsupervised domain adaptation (UDA) for medical image segmentation remains challenging due to substantial domain shifts across imaging modalities,...
By Lalit Maurya, Honghai Liu, Reyer Zwiggelaar
The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and applying posterior‑weighted mean subtraction, followed by a log‑prior correction based on confidence‑weighted predictions. DRC improves cross‑domain accuracy, surpassing zero‑shot CLIP by 4.13 and 5.07 points on ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.
The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
The paper introduces Domain Recentering with Confidence Calibration (DRC), a training‑free technique that adapts CLIP to unlabeled target images by fitting a Gaussian mixture and subtracting a posterior‑weighted average of component means from each embedding. It further corrects residual class bias using a log‑prior adjustment based on confidence‑weighted predictions. DRC outperforms other methods, raising average accuracy on cross‑domain datasets by 4.13 and 5.07 points over zero‑shot CLIP for ViT‑B/16 and ResNet‑50, and maintains gains under ImageNet distribution shifts.
By Youngeun Seol, Jimin Shin, Heeseo Yoon, Uiwon Hwang
arXiv:2507. 18632v2 Announce Type: replace-cross Abstract: Zero-shot domain adaptation is a method for adapting a model to a target domain without utilizing target domain image data.
By Ye-Chan Kim, SeungJu Cha, Si-Woo Kim, Taewhan Kim, Dong-Jin Kim
arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
The paper introduces Concept-Driven Domain Adaptation (CDDA), a three-stage framework that adapts vision‑language models for concept‑to‑example video retrieval in educational settings. CDDA first structures textual embeddings using textbook and teacher‑handbook concept pairs, then transfers this geometry to documentary visuals with a frozen visual encoder, and finally jointly fine‑tunes both encoders with sparse visual concept supervision. On a middle‑school physics benchmark, CDDA outperforms several multimodal baselines in retrieving concept‑driven moments while preserving concrete image‑text alignment.
By Haiming Zhao, Tai Wang, Kun Zhang, Xicheng Peng, Zhiyang Li
Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored.