arXiv Computer Vision By Rushab Rasik Karania, Tomas Maul

Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes

Read the original on arXiv Computer Vision →

The paper introduces the Within-Instance Prototypical Transformer (WIPT), a method that performs single-query test-time prototype adaptation by jointly transforming an unlabelled query with labelled support embeddings to form query-specific class means. Using a frozen ViT-S/16 encoder trained on miniImageNet and evaluated on CUB, EuroSAT, and ISIC, WIPT improves 1‑shot performance on CUB and EuroSAT but not on ISIC, while in 5‑shot settings it outperforms a support‑only Transformer on ISIC but remains weaker than ProtoNet overall. The study also explores multi‑query processing, memory and latency trade‑offs, and analyzes how WIPT alters uncertain versus confident predictions, concluding that the method offers a streaming‑compatible test‑time adaptation that can enhance low‑shot cross‑domain decisions without target‑time optimization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Aug 3

Visual Distribution Anchoring for Efficient Prompt Tuning

arXiv:2607. 28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch.

By Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi
arXiv Machine Learning
Sep 11

Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training Assumptions

The paper critically evaluates common few‑shot learning protocols that rely on pre‑training a model on a large auxiliary set with classes disjoint from the target but drawn from the same visual domain. By comparing no pre‑training, class‑disjoint in‑domain pre‑training, supervised out‑of‑domain pre‑training, and label‑free out‑of‑domain pre‑training across eight datasets and three architectures, the authors find that in‑domain pre‑training yields a 33.41‑point average improvement, while out‑of‑domain pre‑training offers a 23.75‑point gain, revealing a 9.66‑point optimistic bias due to domain overlap. They also demonstrate that a label‑free augmentation strategy can match supervised out‑of‑domain performance and propose a descriptor‑based source‑selection method that closely approximates oracle selection, underscoring the need to move beyond in‑domain pre‑training as the default evaluation protocol.

By Alejandro Galan-Cuenca, Marcelo Saval-Calvo, Antonio Javier Gallego
arXiv Computer Vision
Sep 22

Training-Free Spectral Transductive Refinement for Cross-Domain Few-Shot Classification

The paper introduces Spectral Transductive Refinement (STR), a training‑free method that refines class prototypes at test time using the geometry of a joint k‑nearest‑neighbour graph and a normalized‑Laplacian spectral coordinate system. STR operates solely on frozen visual embeddings, iteratively updating pseudo‑labelled queries to improve one‑shot and few‑shot classification under domain shift. Experiments on ResNet‑18 and ResNet‑10 backbones show STR outperforms single‑prototype baselines and rivals meta‑trained cross‑domain few‑shot methods, achieving the best 1‑shot average across eight target domains.

By Fahim Rahman, S. M. Tanjeeb Meheran Rohan, Md. Taimum Ibne Sayed, Asaduzzaman Herok, Md. Bakhtiar Hasan