arXiv Machine Learning

Advancements in Content-Based Image Retrieval: A Comprehensive Survey of Relevance Feedback Techniques

This survey reviews content‑based image retrieval (CBIR) systems, highlighting their use of visual content for image search and their importance in object detection. It discusses key challenges such as the semantic gap and scalability, and examines relevance feedback (RF) techniques—including long‑term and short‑term learning, weight optimization, and active learning—to iteratively refine search results. The paper also explores machine‑learning and deep‑learning approaches, particularly convolutional neural networks, to improve CBIR accuracy and relevance.

Hugging Face Trending Papers
Jun 3

COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling.

arXiv AI
Sep 1

When Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception

The paper introduces CERES, a closed‑loop multimodal indexing framework that addresses semantic collapse in multimodal generation by building a three‑level semantic pyramid and using scale‑routed cross‑attention to generate images that remain retrievable by their original queries. CERES employs a co‑occurrence‑aware router, a lightweight U‑Net generator, and a soft‑Jaccard coverage objective to ensure generated images cover the intended concepts, verified by re‑indexing with a frozen vision‑language model and an external DINOv2 probe. Experiments on four pansharpening benchmarks show state‑of‑the‑art performance, especially under extreme scale variation, and significant improvements in concept‑query retrieval and image‑text ranking metrics.

By Guangyuan Dong, Chuang Liu, Haoyu Wang, Yangchen Zeng, Jiaqi Zhang, Li Jiuxing, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin, Alexander Lim Han Yang, Yusen Wu
arXiv AI
Aug 24

When Generated Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception

arXiv:2608. 20810v1 Announce Type: cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve.

By Guangyuan Dong, Chuang Liu, Yangchen Zeng, Haoyu Wang, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin
arXiv Computer Vision
Sep 22

Graded-Relevance Composed Multimodal Retrieval for E-commerce Visual Search at Scale

The paper introduces GradCIR, a method for training composed image retrieval (CIR) systems on graded relevance rather than binary relevance. It uses a vision‑language model to generate queries and 4‑level relevance labels, an iterative feedback loop to mine hard negatives, and a hierarchy‑aware angular objective to directly optimize graded labels. Experiments on a Walmart catalog and FashionIQ show significant NDCG improvements and the system is deployed in Walmart’s live visual‑search traffic.

By Anubhav Gupta, Hrushikesh Mohapatra, Prijith Chandra, Asish Mohapatra, Anuj Garg, Arvind Maan, Sudip Datta, Venkat Bulusu, Sitesh Kumar Jalan
arXiv AI
Aug 28

PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

PailitaoGR is a generative image retrieval model that incorporates a latent think-with-images approach to better handle real‑world query images. It uses a target‑focused perception mechanism—comprising a target enhancer and on‑policy distillation—to highlight the search target, and a selective auxiliary‑evidence mechanism—using an auxiliary enhancer and incremental contrastive distillation—to exploit useful side information. Trained on real‑world online image‑search logs, the method achieves an average 13.8 % improvement over existing baselines.

By Xiaomeng Fan, Yueran Liu, Shengyu Zhou, Chenghan Fu, Wanxian Guan, Feng Li, Chuan Yu, Jian Xu, Bo Zheng
arXiv Computation and Language
Aug 27

Recurrence Meets Transformers for Universal Multimodal Retrieval

The paper introduces ReT-2, a unified retrieval model that handles multimodal queries containing both images and text and searches across multimodal document collections. It employs a recurrent Transformer architecture with LSTM-inspired gating to integrate information across layers and modalities, capturing fine-grained visual and textual details. Evaluations on M2KR and M-BEIR benchmarks show state‑of‑the‑art performance, faster inference, and lower memory usage, and the model also boosts downstream tasks in retrieval‑augmented generation pipelines.

By Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Hugging Face Trending Papers
Aug 27

PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

PailitaoGR is a generative image retrieval method that incorporates target-focused perception and selective auxiliary-evidence utilization. It uses a target Enhancer and on-policy distillation to highlight search-target regions, and an auxiliary enhancer with incremental contrastive distillation to exploit auxiliary evidence. Trained on real-world online image-search logs, it achieves an average 13.8% improvement over existing baselines.