arXiv Machine Learning

Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels

arXiv:2604. 06614v2 Announce Type: replace-cross Abstract: Prompt learning has gained significant attention as a parameter-efficient approach for adapting large pre-trained vision-language models to downstream tasks.

arXiv Computer Vision
Sep 16

TecoPrompt: Temporal-Conservative Prompt Learning for Vision-Language Models

TecoPrompt introduces a closed‑loop robust prompt‑learning framework that mitigates label noise in vision‑language models. It uses optimal transport in the CLIP semantic space to generate globally consistent pseudo‑labels, then verifies their reliability by checking trajectory stability over a K‑epoch window and an EMA‑based confidence gate. The verified labels are incorporated into prompt training via a tri‑group objective, yielding significant accuracy gains across multiple noisy datasets, such as a 0.843 accuracy on OxfordPets with 50% asymmetric noise.

By Zeyi Shao, Haowen Hua, Jiaxin Zhang, John See, Zeyd Boukhers, Cong Yang
arXiv Computer Vision
6d ago

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.

By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
arXiv Computer Vision
Aug 25

Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label

The paper introduces Language-driven Dense Semantic Adaptor (LDSA) for multi-label image classification with incomplete annotations. LDSA leverages multimodal pretrained CLIP models to extract prior-adaptive relationships, employing a densely contrastive adaptor for visual contrastive constraints and a language-driven interactive decoder with class-specific prompt tuning. Experiments show LDSA achieves state‑of‑the‑art performance on public benchmarks and reveals implicit semantic relationships through its learning scheme.

By Cheng Chen, Yifan Zhao, Jia Li
arXiv AI
Sep 2

Guided Prompt Evolution for Vision-Language Models Adaptation

The paper introduces EvoPrompt, a framework for adapting vision‑language models to new tasks with limited data while preventing catastrophic forgetting. EvoPrompt uses a Modality‑Shared Prompt Projector to create hierarchical prompts and an evolutionary training strategy that separates low‑rank updates into directional and magnitude components, preserving learned semantic directions. Experiments show that EvoPrompt achieves state‑of‑the‑art few‑shot performance while maintaining the original zero‑shot capabilities of the pre‑trained models.

By Enming Zhang, Jiayang Li, Yanlong Wang, Yanru Wu, Zhenyu Liu, Yang Li