Preserve-and-Compose Training for Composed Image Retrieval
Read the original on arXiv Computer Vision →The paper introduces Preserve-and-Compose Training (PACT) for composed image retrieval, a task where a query image is modified by a textual instruction while preserving visual content from a reference image. PACT learns from image–text–text triplets, using target captions for supervision and visual evidence from the source image to maintain relevant details, without requiring target images or gallery updates. The authors also propose Chord scoring, which blends target similarity with source-relative directional agreement in a frozen image space, and demonstrate that this combined approach yields strong retrieval performance across multiple zero-shot CIR benchmarks and various backbones.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.