arXiv:2609.14100v1 Announce Type: cross
Abstract: Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image...
By Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra
The paper introduces GradCIR, a method for training composed image retrieval (CIR) systems on graded relevance rather than binary relevance. It uses a vision‑language model to generate queries and 4‑level relevance labels, an iterative feedback loop to mine hard negatives, and a hierarchy‑aware angular objective to directly optimize graded labels. Experiments on a Walmart catalog and FashionIQ show significant NDCG improvements and the system is deployed in Walmart’s live visual‑search traffic.
By Anubhav Gupta, Hrushikesh Mohapatra, Prijith Chandra, Asish Mohapatra, Anuj Garg, Arvind Maan, Sudip Datta, Venkat Bulusu, Sitesh Kumar Jalan
arXiv:2607. 22919v1 Announce Type: cross Abstract: Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification.
By Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani, Mubarak Shah, Kushal Kafle
MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.
By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang
Visual search on large e-commerce catalogs must serve both "similarity" queries that ask for items resembling an uploaded image and "modifier" queries that comprise an image and text describing a desi...
Adapting CLIP for zero-shot sketch-based image retrieval (ZS-SBIR) via prompt learning faces a fundamental tension: the model must bridge the sketch-photo domain gap through task-specific adaptation, yet the added flexibility risks overfitting to seen training categories and eroding CLIP's zero-shot generalization. We present SeCo-SBIR, a semantically consistent prompt learning framework that resolves this tension from both sides.
arXiv:2608. 12987v1 Announce Type: cross Abstract: Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly.
By Kaipeng Li, Haitao Yu, Xuanchen Zhou
arXiv:2607. 18695v1 Announce Type: cross Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors.
By Gautam Rajendrakumar Gare, Jia Shi, Zhiqiu Lin, Deepak Pathak, John Galeotti, Deva Ramanan
arXiv:2607. 29000v1 Announce Type: cross Abstract: Multimodal information can improve the accuracy of click-through rate (CTR) prediction and effectively alleviate item cold-start and long-tail problems.
By Huanyu Liu, Baining Chen, Hui Liu, Zengyang Li, Ziyi Huang
arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.
By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja
arXiv:2608.20886v1 Announce Type: cross
Abstract: Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify,...
By Enjun Du, Siyi Liu, Zirong Chen, Xinyu Zuo, Jinwen Luo, Ruiwen Tao, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang
The paper introduces SEAL, a plug‑and‑play module that enhances single‑image sticker personalization in diffusion models by adding a semantic‑guided spatial attention loss, a split‑merge token strategy, and structure‑aware layer restriction. SEAL integrates without altering the U‑Net backbone and addresses overfitting issues such as visual entanglement and structural rigidity. Alongside SEAL, the authors release StickerBench, a large sticker dataset with six attribute tags to enable systematic evaluation of identity preservation and contextual controllability.
By Changhyun Roh, Yonghyun Jeong, Jonghyun Lee, Chanho Eom, Jihyong Oh