arXiv AI By Dong Liu, Sreyashi Nag

Query Brand Entity Linking in E-Commerce Search

Read the original on arXiv AI →

The paper "Query Brand Entity Linking in E-Commerce Search" addresses the challenge of matching short, unstructured user search queries to the correct brand entities in large e‑commerce catalogs. It proposes two scalable solutions: a cascaded pipeline that first detects brand mentions and then disambiguates them, and a single‑stage extreme multiclass classifier that directly maps queries to brand identifiers. Extensive multilingual evaluation and an online experiment show that both methods significantly improve brand recall while preserving high precision, resulting in measurable gains in customer engagement.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 2

Semantic Retrieval for Product Search in E-Commerce

arXiv:2606. 01504v1 Announce Type: cross Abstract: Semantic retrieval in e-commerce must handle short, noisy, and colloquial queries over large product catalogs with fine-grained attribute distinctions.

By Nikhil Kothari, Saksham Samdani, Ritam Mallick, Praveen Gupta, Ankit Vijay, Surender Kumar
arXiv Machine Learning
Jul 28

SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads

arXiv:2607. 23121v1 Announce Type: cross Abstract: Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories).

By Congfei Zhang, Jingxiao Ma, Xiaodong Liu, Hsiang-wei Chao, Siman Wang, Ge Liu, Shantanu Aggarwal, Vincent Zhang, Meghana Missula, Rachel Liao, Zichu Li, Xiao Bai, Yunzhi Zhou, Yajun Wang, Zhe Liu, Jinchao Li, Yu Zhang
arXiv Machine Learning
Aug 31

Mine and Refine: Optimizing Graded Relevance in E-commerce Semantic Search Retrieval

The paper introduces Mine and Refine, a two‑stage contrastive training framework designed to improve embedding‑based retrieval for large‑scale e‑commerce search. It tackles graded relevance, hard‑sample mining, and unstable similarity separability by using a lightweight LLM as a scalable labeler and a multi‑level circle loss to enforce margin‑controlled separation across relevance levels. The method has been deployed in production across multiple product verticals, yielding statistically significant increases in user engagement, gross order value, and retrieval relevance metrics.

By Jiaqi Xi, Raghav Saboo, Luming Chen, Johny Rufus, Aditya Dodda, Ved Sampath, Kenny Chi, Elyse Winer, Akshad Viswanathan, Martin Wang, Sudeep Das
arXiv AI
Jun 24

AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

arXiv:2606. 24655v1 Announce Type: cross Abstract: The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction.

By Murilo Gazzola, Hugo Gobato Souto, Samuel Silva, J\'ulia Schubert Peixoto, Felipe Siqueira, Andr\'e Luis Pedroso de Morais, Caio Gomes
arXiv Machine Learning
5d ago

Retail Product Search: A Practical Approach at Target

The paper describes a hybrid search system developed at Target that combines lexical and vector search to improve retail product search. It details data processing, embedding training, precision control, multi‑channel result fusion—specifically weighted interleaving—and performance optimizations for low latency. The system achieved measurable gains in click‑through rate, order conversion, and demand per visitor while reducing zero‑result searches.

By Darshan Sonagara, Qujiaheng Zhang, Ankit Singh, Alex Li