arXiv Machine Learning

SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads

arXiv:2607. 23121v1 Announce Type: cross Abstract: Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories).

Hugging Face Trending Papers
Jul 29

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items.

arXiv Machine Learning
5d ago

Retail Product Search: A Practical Approach at Target

The paper describes a hybrid search system developed at Target that combines lexical and vector search to improve retail product search. It details data processing, embedding training, precision control, multi‑channel result fusion—specifically weighted interleaving—and performance optimizations for low latency. The system achieved measurable gains in click‑through rate, order conversion, and demand per visitor while reducing zero‑result searches.

By Darshan Sonagara, Qujiaheng Zhang, Ankit Singh, Alex Li
arXiv AI
Aug 24

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

The paper proposes a single hierarchical Semantic ID (SID) system to unify product identification across multiple merchants in e-commerce. By learning SID representations from product content, the authors demonstrate that ranking algorithms can aggregate consumer affinity and product performance over SID prefixes, improving offline relevance and online engagement. For query reformulation, SID concepts guide navigation and refinement, yielding better intent preservation and higher-quality suggestions compared to taxonomy or raw query transitions.

By Steven Xu, Sanjyot Thete, Saathvik Dirisala, Raghav Saboo, Nimesh Sinha, Leo Shao, Elyse Winer, Sudeep Das, Martin Wang, Kyle MacDonald
arXiv Machine Learning
Aug 31

Mine and Refine: Optimizing Graded Relevance in E-commerce Semantic Search Retrieval

The paper introduces Mine and Refine, a two‑stage contrastive training framework designed to improve embedding‑based retrieval for large‑scale e‑commerce search. It tackles graded relevance, hard‑sample mining, and unstable similarity separability by using a lightweight LLM as a scalable labeler and a multi‑level circle loss to enforce margin‑controlled separation across relevance levels. The method has been deployed in production across multiple product verticals, yielding statistically significant increases in user engagement, gross order value, and retrieval relevance metrics.

By Jiaqi Xi, Raghav Saboo, Luming Chen, Johny Rufus, Aditya Dodda, Ved Sampath, Kenny Chi, Elyse Winer, Akshad Viswanathan, Martin Wang, Sudeep Das
arXiv Machine Learning
Jun 30

Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation

arXiv:2606. 29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regimes.

By Zhe Dong (University of Maine at Presque Isle), Fang Qin (Stanford University), Manish Shah (Independent Researcher), Yicheng Wang (Independent Researcher)