Rethinking Text-Based Image Retrieval in Specific Domain
arXiv:2608. 10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress.
This survey reviews content‑based image retrieval (CBIR) systems, highlighting their use of visual content for image search and their importance in object detection. It discusses key challenges such as the semantic gap and scalability, and examines relevance feedback (RF) techniques—including long‑term and short‑term learning, weight optimization, and active learning—to iteratively refine search results. The paper also explores machine‑learning and deep‑learning approaches, particularly convolutional neural networks, to improve CBIR accuracy and relevance.
arXiv:2608. 10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress.
Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling.
The paper introduces CERES, a closed‑loop multimodal indexing framework that addresses semantic collapse in multimodal generation by building a three‑level semantic pyramid and using scale‑routed cross‑attention to generate images that remain retrievable by their original queries. CERES employs a co‑occurrence‑aware router, a lightweight U‑Net generator, and a soft‑Jaccard coverage objective to ensure generated images cover the intended concepts, verified by re‑indexing with a frozen vision‑language model and an external DINOv2 probe. Experiments on four pansharpening benchmarks show state‑of‑the‑art performance, especially under extreme scale variation, and significant improvements in concept‑query retrieval and image‑text ranking metrics.
arXiv:2608. 20810v1 Announce Type: cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve.
The paper introduces GradCIR, a method for training composed image retrieval (CIR) systems on graded relevance rather than binary relevance. It uses a vision‑language model to generate queries and 4‑level relevance labels, an iterative feedback loop to mine hard negatives, and a hierarchy‑aware angular objective to directly optimize graded labels. Experiments on a Walmart catalog and FashionIQ show significant NDCG improvements and the system is deployed in Walmart’s live visual‑search traffic.
arXiv:2609.01456v1 Announce Type: cross Abstract: Composed image retrieval (CIR) retrieves a target image from a reference image and a text modification. This paper studies metadata-available CIR rer...
PailitaoGR is a generative image retrieval model that incorporates a latent think-with-images approach to better handle real‑world query images. It uses a target‑focused perception mechanism—comprising a target enhancer and on‑policy distillation—to highlight the search target, and a selective auxiliary‑evidence mechanism—using an auxiliary enhancer and incremental contrastive distillation—to exploit useful side information. Trained on real‑world online image‑search logs, the method achieves an average 13.8 % improvement over existing baselines.
Visual search on large e-commerce catalogs must serve both "similarity" queries that ask for items resembling an uploaded image and "modifier" queries that comprise an image and text describing a desi...
arXiv:2608.23102v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is an emerging paradigm in content-based image retrieval that enables users to formulate compositional queries by combin...
The paper introduces ReT-2, a unified retrieval model that handles multimodal queries containing both images and text and searches across multimodal document collections. It employs a recurrent Transformer architecture with LSTM-inspired gating to integrate information across layers and modalities, capturing fine-grained visual and textual details. Evaluations on M2KR and M-BEIR benchmarks show state‑of‑the‑art performance, faster inference, and lower memory usage, and the model also boosts downstream tasks in retrieval‑augmented generation pipelines.
PailitaoGR is a generative image retrieval method that incorporates target-focused perception and selective auxiliary-evidence utilization. It uses a target Enhancer and on-policy distillation to highlight search-target regions, and an auxiliary enhancer with incremental contrastive distillation to exploit auxiliary evidence. Trained on real-world online image-search logs, it achieves an average 13.8% improvement over existing baselines.
arXiv:2609.08188v1 Announce Type: new Abstract: Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retriever...