Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image from a multimodal query consisting of a reference image and an edit text describing the desired modification. Recent ZS-CIR studies have relied on projection-based methods that map a reference image into pseudo-word tokens in the text embedding space.
Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image by editing a reference image with a natural-language instruction, without relying on domain-specific annotated triplets. Most existing ZS-CIR methods rely on textual inversion to translate the reference image into pseudo-text tokens and then compose them with the instruction via simple concatenation in the text space, which can be lossy and brittle for fine-grained semantics.
arXiv:2607. 00374v1 Announce Type: cross Abstract: Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification.
By Jingjing Zhang, Lei Zhang, Zheren Fu, Zhendong Mao
arXiv:2609.10008v1 Announce Type: new
Abstract: Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gal...
By Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar, Abdelrahman Mohamed Shaker, Rao Anwer
The paper introduces Preserve-and-Compose Training (PACT) for composed image retrieval, a task where a query image is modified by a textual instruction while preserving visual content from a reference image. PACT learns from image–text–text triplets, using target captions for supervision and visual evidence from the source image to maintain relevant details, without requiring target images or gallery updates. The authors also propose Chord scoring, which blends target similarity with source-relative directional agreement in a frozen image space, and demonstrate that this combined approach yields strong retrieval performance across multiple zero-shot CIR benchmarks and various backbones.
By Sehyun Kwon
arXiv:2608.30584v1 Announce Type: new
Abstract: Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on si...
By Xingjian Wang, Shijian Wang, Yibo Wang, Zihao Yu, Runhao Fu, Xuelian Cheng, Zongyuan Ge
The paper introduces CLEAR, a CLoze-style rEAsoning-based Re-ranking framework for Compositional Zero-Shot Learning. CLEAR treats primitive variations as context-driven activations of concrete visual cues rather than independent entities, extracting conditional variants in a coarse-to-fine manner and performing cloze-style reasoning to infer high-level semantics. Experiments show that CLEAR consistently improves base models and surpasses state-of-the-art methods on the C-GQA and MIT-States datasets.
By Weize Li, Zhicheng Zhao, Fei Su
arXiv:2607. 07117v1 Announce Type: cross Abstract: In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image.
By Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee
arXiv:2609.12965v1 Announce Type: new
Abstract: Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Mo...
By Mang Ye, Yucheng Ji, Yang Bai, Min Cao, Siyuan Chai, Bo Du, Min Zhang
MulVec is a training‑free zero‑shot composed image retrieval method that matches a target image to a gallery using a reference image and a text edit. It introduces a role‑aware query system that separates the target description into four retrieval roles—Global, Desired, Preserve, and Forbidden—each mapped to specific probe vectors. By combining global and local visual representations, MulVec achieves significant performance gains on CIRCO, CIRR, and FashionIQ datasets, improving CIRCO mAP@5 by up to 23.0% over prior methods.
By Zihao Zhang, Dayan Wu, Xinze Liu, Hengjie Zhu, Yiliang Zhu, Ding Wang, Peng Fu, Zheng Lin, Weiping Wang
The paper introduces ExpArt-KG, a knowledge graph tailored to the artwork domain, and a retrieval‑augmented generation framework that alternates between generating answers and retrieving relevant facts from the graph. By using a correctness judgment to guide the search, the method efficiently gathers the necessary factual information, improving the detail of image explanations while reducing external knowledge retrieval costs. Experimental results demonstrate that the approach maintains generation quality comparable to fixed‑iteration methods.
By Yuta Kato, Shintaro Ozaki, Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe
arXiv:2609.37407v1 Announce Type: new
Abstract: While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a ma...
By Xianghan Wei, Xiaoda Yang, Zhi Wang, An Pan, Daoan Zhang, Huayi Zhang, Yan Zhang, Wei Xu, Zishun Liao, Jianwen Lou