MMRM: A Multiplex Multimodal Representation Model for Product Ranking in E-commerce Search
arXiv:2607. 11030v1 Announce Type: cross Abstract: Multimodal information is pivotal for e-commerce search ranking.
arXiv:2607. 17499v1 Announce Type: new Abstract: The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and mixed-intent instructions.
arXiv:2607. 11030v1 Announce Type: cross Abstract: Multimodal information is pivotal for e-commerce search ranking.
arXiv:2607. 29002v1 Announce Type: new Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone.
arXiv:2605.01278v3 Announce Type: replace Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified unders...
arXiv:2607. 29213v1 Announce Type: cross Abstract: Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent.
arXiv:2607. 03886v1 Announce Type: cross Abstract: Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align with the users' search queries.
arXiv:2511. 12449v3 Announce Type: replace-cross Abstract: Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding.
The paper introduces GradCIR, a method for training composed image retrieval (CIR) systems on graded relevance rather than binary relevance. It uses a vision‑language model to generate queries and 4‑level relevance labels, an iterative feedback loop to mine hard negatives, and a hierarchy‑aware angular objective to directly optimize graded labels. Experiments on a Walmart catalog and FashionIQ show significant NDCG improvements and the system is deployed in Walmart’s live visual‑search traffic.
arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
The paper describes a hybrid search system developed at Target that combines lexical and vector search to improve retail product search. It details data processing, embedding training, precision control, multi‑channel result fusion—specifically weighted interleaving—and performance optimizations for low latency. The system achieved measurable gains in click‑through rate, order conversion, and demand per visitor while reducing zero‑result searches.
Visual search on large e-commerce catalogs must serve both "similarity" queries that ask for items resembling an uploaded image and "modifier" queries that comprise an image and text describing a desi...
arXiv:2608.24053v1 Announce Type: new Abstract: Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space...
arXiv:2607. 27172v1 Announce Type: cross Abstract: Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall.