arXiv Machine Learning By Seokhee Jin, Changhwan Sung, Sunung Mun, Hoyoung Kim, Jungseul Ok

AdaBoosting Text Prompts for Vision-Language Models

Read the original on arXiv Machine Learning →

arXiv:2607. 00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
6d ago

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.

By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
arXiv Computer Vision
Aug 26

What Does Prompt Learning Change? -A Natural-Language Concept Analysis of Vision-Language Models

Prompt learning modifies vision‑language models by optimizing continuous prompt vectors, yet the resulting prompts are hard to interpret in natural language. PromptSpLiCE is a post‑hoc method that rewrites each class‑conditioned text embedding as a sparse mix of concepts from a fixed dictionary, enabling a direct comparison of concept profiles before and after prompt learning. Across 11 image‑classification datasets, the method shows that only about 1.6 of the initial top‑10 concepts remain after learning, and that larger profile changes correlate with higher accuracy gains, while a derived gradient expression offers geometric insight into loss sensitivity.

By Ryo Kamiya, Hiroshi Kera, Kazuhiko Kawamoto
arXiv AI
Sep 2

Guided Prompt Evolution for Vision-Language Models Adaptation

The paper introduces EvoPrompt, a framework for adapting vision‑language models to new tasks with limited data while preventing catastrophic forgetting. EvoPrompt uses a Modality‑Shared Prompt Projector to create hierarchical prompts and an evolutionary training strategy that separates low‑rank updates into directional and magnitude components, preserving learned semantic directions. Experiments show that EvoPrompt achieves state‑of‑the‑art few‑shot performance while maintaining the original zero‑shot capabilities of the pre‑trained models.

By Enming Zhang, Jiayang Li, Yanlong Wang, Yanru Wu, Zhenyu Liu, Yang Li
arXiv Machine Learning
Jul 9

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

arXiv:2607. 07179v1 Announce Type: cross Abstract: Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents.

By Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia