Hugging Face Trending Papers

SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval

Adapting CLIP for zero-shot sketch-based image retrieval (ZS-SBIR) via prompt learning faces a fundamental tension: the model must bridge the sketch-photo domain gap through task-specific adaptation, yet the added flexibility risks overfitting to seen training categories and eroding CLIP's zero-shot generalization. We present SeCo-SBIR, a semantically consistent prompt learning framework that resolves this tension from both sides.

arXiv Computer Vision
6d ago

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.

By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
arXiv AI
6d ago

Prompt-Based Continual Compositional Zero-Shot Learning

The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.

By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv Computer Vision
Aug 28

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

The paper introduces G2D, a training‑free framework that combines a discriminative model (CLIP) for broad candidate retrieval with a generative vision‑language model for fine‑grained, image‑grounded verification. By using CLIP’s top‑K shortlist and a structured prior from candidate names and probabilities, G2D focuses generative reasoning on uncertain samples, achieving an average accuracy of 68.85% across eight benchmarks—higher than both CLIP alone (59.35%) and the standalone generative model (63.11%). The approach also adapts to various generator configurations and extends to other models such as DCLIP, WaffleCLIP, and CuPL.

By Zehua Hao, Fang Liu, Qinliang Wang, Yaoyang Du, Xinyan Huang, Puhua Chen
Hugging Face Trending Papers
Aug 27

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is a training‑free framework that combines a discriminative model (CLIP) for broad candidate retrieval with a generative vision‑language model for fine‑grained, image‑grounded verification. By using CLIP’s top‑K shortlist and a structured prior from candidate names and probabilities, G2D focuses generative reasoning on uncertain samples, employing fixed confidence routing, entropy‑adaptive candidate sizing, and trie‑constrained decoding to produce a single valid output. Across eight benchmarks, G2D achieves an average accuracy of 68.85%, outperforming both CLIP (59.35%) and the standalone generative model (63.11%), and it also transfers effectively to other models such as DCLIP, WaffleCLIP, and CuPL.

arXiv Computer Vision
Aug 25

SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space

SketchFlow is a new generative framework for creating high‑quality vector sketches from text prompts. It uses a Gaussian Mixture Model prior in the CLIP latent space and an Optimal Transport Conditional Flow Matching model to map this prior to sketch features, which are then decoded by a Hybrid Diffusion Decoder combining 1D U‑Net and Transformer architectures. The approach achieves superior visual quality and human‑like drawing styles, and supports zero‑shot synthesis for unseen concepts and smooth semantic interpolation.

By Jin Zhou, Hongliang Yang, Pengfei Xu, Hui Huang
Hugging Face Trending Papers
Jul 30

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR).

arXiv Computer Vision
Sep 25

Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

The paper introduces MK‑FSS, a few‑shot segmentation framework that leverages Multimodal Large Language Models (MLLMs) to extract spatial and semantic target knowledge from query images. Spatial knowledge is encoded into a memory representation and fused with support‑guided memory via a dual‑memory debate‑fusion module, while semantic knowledge is turned into a textual feature and combined with multi‑scale query features through a progressive cross‑modal prompt generator. Together, these components produce a robust target representation that improves segmentation performance over existing methods.

By Yijun Hu, Heng Fan, Libo Zhang
Hugging Face Trending Papers
Jul 15

Fine-grained CLIP fine-tuning with self-annotated region alignment

Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-training a vision-language model from scratch, a series of works aim to enhance the fine-grained ability of CLIP through a fine-tuning scheme.