arXiv:2607. 28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch.
By Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi
The paper critically evaluates common few‑shot learning protocols that rely on pre‑training a model on a large auxiliary set with classes disjoint from the target but drawn from the same visual domain. By comparing no pre‑training, class‑disjoint in‑domain pre‑training, supervised out‑of‑domain pre‑training, and label‑free out‑of‑domain pre‑training across eight datasets and three architectures, the authors find that in‑domain pre‑training yields a 33.41‑point average improvement, while out‑of‑domain pre‑training offers a 23.75‑point gain, revealing a 9.66‑point optimistic bias due to domain overlap. They also demonstrate that a label‑free augmentation strategy can match supervised out‑of‑domain performance and propose a descriptor‑based source‑selection method that closely approximates oracle selection, underscoring the need to move beyond in‑domain pre‑training as the default evaluation protocol.
By Alejandro Galan-Cuenca, Marcelo Saval-Calvo, Antonio Javier Gallego
The paper introduces Spectral Transductive Refinement (STR), a training‑free method that refines class prototypes at test time using the geometry of a joint k‑nearest‑neighbour graph and a normalized‑Laplacian spectral coordinate system. STR operates solely on frozen visual embeddings, iteratively updating pseudo‑labelled queries to improve one‑shot and few‑shot classification under domain shift. Experiments on ResNet‑18 and ResNet‑10 backbones show STR outperforms single‑prototype baselines and rivals meta‑trained cross‑domain few‑shot methods, achieving the best 1‑shot average across eight target domains.
By Fahim Rahman, S. M. Tanjeeb Meheran Rohan, Md. Taimum Ibne Sayed, Asaduzzaman Herok, Md. Bakhtiar Hasan
arXiv:2609.22323v1 Announce Type: cross
Abstract: Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required...
By Neeraj Yadav
arXiv:2607. 17467v1 Announce Type: cross Abstract: Few-shot Test-Time Domain Adaptation (FSTT-DA) seeks to adapt models to novel domains using only a handful of unlabeled target samples.
By Siobhan Reid, Zhixiang Chi, Li Gu, Omid Reza Heidari, Ziqiang Wang, Yang Wang
arXiv:2608.29395v1 Announce Type: new
Abstract: Vision-language models such as CLIP and SigLIP provide strong zero-shot recognition, but their predictions can degrade when deployed on target data tha...
By Pedram MohajerAnsari, Amir Salarpour, Run Wang, Mert D. Pes\'e
ReVisIT is a train‑free framework that turns retrieved image‑label pairs into units of visual thought, combining structured class definitions, multimodal retrieval, and alternating user/assistant injection before joint decoding. On several benchmarks—including Fast Open MiniImageNet, Bongard‑OpenWorld, and the newly released MAAC‑Bench—ReVisIT achieves performance comparable to or surpassing large, trained models while using far fewer parameters. The approach demonstrates that high‑quality retrieval and a simple turns layer can provide a universal performance boost across diverse multimodal tasks.
By Bingchen Huang, Zhiling Wang, Yifu Chen, Yuanchao Du
The paper investigates whether language prompts selected by zero‑shot accuracy remain effective after visual adaptation in source‑free cross‑domain few‑shot learning. Using a paired protocol, the authors compare generic class‑name templates with detailed class descriptions before and after Low‑Rank Adaptation (LoRA) on datasets such as EuroSAT, CropDisease, ISIC, and ChestX. They identify two regimes: semantic saturation, where detailed prompts are already useful before adaptation, and semantic emergence, where detailed prompts become more useful only after visual representation updates, driven by changes in prediction patterns.
By Wei Liu, Xing Deng, Haijian Shao
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
By Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang
arXiv:2604.15678v2 Announce Type: replace
Abstract: Pretrained Vision-Language Models (VLMs) like CLIP show promise in continual learning, but existing Few-Shot Class-Incremental Learning (FSCIL) met...
By Eunju Lee, MiHyeon Kim, JuneHyoung Kwon, Yoonji Lee, JiHyun Kim, Soojin Jang, YoungBin Kim
The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.
By Kentaro Oda
The paper demonstrates that sharing a deep encoder alone does not eliminate the confounding effects in task-comparison scores. By introducing a conditional two‑discriminator discrepancy within the embedding space, the authors achieve robust detection of task changes, maintaining stability under input rotations and accurately tracking label‑permutation drift. This approach, integrated into a mixture‑of‑heads framework, outperforms traditional novelty triggers and generalizes across multiple backbones and datasets, including ImageNet‑21k ViT‑B/16, DINOv2, and CIFAR‑100.
By Kentaro Oda