arXiv Machine Learning By Liangzhi Li, Bowen Wang, Yiming Qian, Thorsten Neumann, Xia Xie, Guangshun Li

The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

Read the original on arXiv Machine Learning →

The paper investigates the blending ratio used in few‑shot adaptation of vision‑language models, which combines a zero‑shot text prototype with the mean of labeled image features. It shows that the theoretically optimal ratio—derived from a closed‑form mean‑squared error minimizer—does not align with the ratio that actually maximizes performance, falling short by an average of 8.5 points. Moreover, a leave‑one‑out estimate on the support set achieves near‑oracle performance, and validation‑free linear probes outperform even oracle‑tuned blends, indicating that the hyperparameter can be set near‑optimally without external validation data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 6

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

arXiv:2608. 01207v2 Announce Type: replace-cross Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poorly to vision-language models (VLMs): recent work shows that simple majority voting beats selection methods built on the model's own self-verification, apparently because at the selection layer an image-grounded answer and a confident guess from the language prior look the same.

By Puzhuo Zheng, Hasan Kurban
arXiv Computer Vision
Sep 2

Can Scene Text Recognition Read Rare Compositions?

The paper reports that scene text recognition models, while achieving 89–97% accuracy on standard benchmarks, perform significantly worse on rare word–trigram combinations, with a 10–18 point drop in accuracy at the rare‑word/rare‑trigram corner across multiple languages and models. Scaling the vision backbone improves overall accuracy but does not alleviate this corner‑specific deficit. The authors identify the autoregressive decoder’s lexical prior as the root cause and show that architectural changes—specifically moving from autoregressive to CTC decoding—yield the largest improvement for these rare compositions.

By Genpei Zhang