The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models
Read the original on arXiv Machine Learning →The paper investigates the blending ratio used in few‑shot adaptation of vision‑language models, which combines a zero‑shot text prototype with the mean of labeled image features. It shows that the theoretically optimal ratio—derived from a closed‑form mean‑squared error minimizer—does not align with the ratio that actually maximizes performance, falling short by an average of 8.5 points. Moreover, a leave‑one‑out estimate on the support set achieves near‑oracle performance, and validation‑free linear probes outperform even oracle‑tuned blends, indicating that the hyperparameter can be set near‑optimally without external validation data.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.