Hugging Face Trending Papers

Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster

arXiv Machine Learning
Aug 31

Mine and Refine: Optimizing Graded Relevance in E-commerce Semantic Search Retrieval

The paper introduces Mine and Refine, a two‑stage contrastive training framework designed to improve embedding‑based retrieval for large‑scale e‑commerce search. It tackles graded relevance, hard‑sample mining, and unstable similarity separability by using a lightweight LLM as a scalable labeler and a multi‑level circle loss to enforce margin‑controlled separation across relevance levels. The method has been deployed in production across multiple product verticals, yielding statistically significant increases in user engagement, gross order value, and retrieval relevance metrics.

By Jiaqi Xi, Raghav Saboo, Luming Chen, Johny Rufus, Aditya Dodda, Ved Sampath, Kenny Chi, Elyse Winer, Akshad Viswanathan, Martin Wang, Sudeep Das
arXiv Machine Learning
Jun 2

Semantic Retrieval for Product Search in E-Commerce

arXiv:2606. 01504v1 Announce Type: cross Abstract: Semantic retrieval in e-commerce must handle short, noisy, and colloquial queries over large product catalogs with fine-grained attribute distinctions.

By Nikhil Kothari, Saksham Samdani, Ritam Mallick, Praveen Gupta, Ankit Vijay, Surender Kumar
arXiv AI
Sep 25

Cross-Country Code-Mixing for Generative Recommendation

Cross-Country Code-Mixing for Generative Recommendation (CMRec) is a framework that enhances generative recommendation across different countries by injecting cross-country supervision at the data level. It learns a shared semantic codebook from multi-modal content and behavioral co-occurrence, then synthesizes mixed-country sequences through token-level substitutions that respect both static and dynamic constraints. A context-aware loss reweights these mixed samples based on their plausibility, leading to improved recommendation quality in data-sparse countries while maintaining performance in data-rich markets, as demonstrated by significant gains in advertising revenue and orders in real-world e-commerce experiments.

By Yuan Gao, Hao Deng, Haibo Xing, Yi Xu, Lingyu Mu, Jinxin Hu, Yu Zhang, Xiaoyi Zeng
arXiv Machine Learning
Jul 28

SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads

arXiv:2607. 23121v1 Announce Type: cross Abstract: Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories).

By Congfei Zhang, Jingxiao Ma, Xiaodong Liu, Hsiang-wei Chao, Siman Wang, Ge Liu, Shantanu Aggarwal, Vincent Zhang, Meghana Missula, Rachel Liao, Zichu Li, Xiao Bai, Yunzhi Zhou, Yajun Wang, Zhe Liu, Jinchao Li, Yu Zhang
arXiv Machine Learning
1d ago

RPTune: Learned Context Curation for LLM Catalog Search

RPTune is an end‑to‑end framework that improves in‑context catalog search for small merchant businesses by learning to curate product catalogs and fine‑tuning large language models (LLMs) with catalog‑grounded supervision. It uses an encoder‑reorganizer curator to order and prune products based on LLM feedback, and then applies context‑relative rewards during LLM post‑training. Across seven real merchants and 100 complex conversational queries per merchant, RPTune boosts search accuracy by up to 31.4 percentage points from curation alone and an additional 10.3 points on average from post‑training.

By Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan \"O. Ar{\i}k
arXiv AI
Sep 10

Exploring Bottom-Up Clustering for Creating Semantic IDs

The paper proposes a new algorithm for generating Semantic IDs that are both unique and preserve the structure of the original embedding space. By employing bottom‑up clustering, the method maintains local structure, leading to higher clustering quality. This improved structure enhances the utility of the Semantic IDs for downstream generative retrieval tasks.

By Leah Woldemariam, Sudhanshu Garg, Taha Belkhouja, Charles Kim-Yip, Ali Sahami
arXiv AI
Aug 24

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

The paper proposes a single hierarchical Semantic ID (SID) system to unify product identification across multiple merchants in e-commerce. By learning SID representations from product content, the authors demonstrate that ranking algorithms can aggregate consumer affinity and product performance over SID prefixes, improving offline relevance and online engagement. For query reformulation, SID concepts guide navigation and refinement, yielding better intent preservation and higher-quality suggestions compared to taxonomy or raw query transitions.

By Steven Xu, Sanjyot Thete, Saathvik Dirisala, Raghav Saboo, Nimesh Sinha, Leo Shao, Elyse Winer, Sudeep Das, Martin Wang, Kyle MacDonald
arXiv AI
Aug 26

Tlow: Flow-based Item Tokenizer for Recommendation

The paper introduces Tlow, a flow-based item tokenizer that transforms raw semantic embeddings into a latent space following a standard normal distribution, enabling independent tokenization and simplifying distributional complexity. Tlow incorporates codebook guidance to align token embeddings with the codebook space, producing semantically clear token IDs. Experiments on four public datasets and an online multi‑modal retrieval task on WeChat show that Tlow improves recommendation performance and increases user click‑through rates by over 10%.

By Nian Li, Chonggang Song, Jingtao Ding, Lingling Yi, Yong Li, Qingmin Liao