arXiv AI

Beyond Co-purchase Relation: Evolution of Complementary Recommendations at Allegro

The paper introduces AlleCompanion, a large‑scale retrieval framework for complementary product recommendations at Allegro.com. It addresses the challenge of noisy co‑purchase data by combining data‑level filtering, a category‑constrained Two Tower architecture, and a multi‑source Complementary Categories Mapping (ComCat) that incorporates expert rules, human feedback, LLM reasoning, and statistical mining. Experiments show that these explicit category constraints and neural models effectively reduce noise, improving recommendation relevance and driving significant GMV growth for both organic discovery and sponsored placements.

Hugging Face Trending Papers
Sep 4

Beyond Co-purchase Relation: Evolution of Complementary Recommendations at Allegro

The paper introduces AlleCompanion, a large‑scale retrieval framework used by Allegro.com to improve complementary product recommendations. It tackles the problem of noisy co‑purchase data by applying data‑level filtering, a category‑constrained Two Tower architecture, and a Category Adapter that limits candidates to logically complementary categories. The system also incorporates a multi‑source Complementary Categories Mapping (ComCat) that blends expert rules, human feedback, LLM reasoning, and statistical mining to refine recommendations, resulting in higher GMV for organic discovery and increased revenue from sponsored placements.

arXiv Machine Learning
Sep 10

A Multi-Source Ensemble Approach to Candidate Generation for Alternative Vacation Rental Property Recommendations

The paper studies candidate generation for alternative vacation rental recommendations, comparing collaborative filtering, shallow embeddings, and graph neural network (GNN) methods on a platform with over 2 million active properties. A hybrid model that combines item-based collaborative filtering with GNN-based retrieval achieves a 14.8% higher Recall@300 than the best baseline, leveraging each method’s strengths: collaborative filtering for well-interacted properties and GNNs for diverse, cold-start alternatives. The authors also show that stronger candidate pools improve downstream ranking quality, though the exact impact is intertwined with ranker training.

By Syed Mohammed Arshad Zaidi, Eric Rincon, Shayan Hassantabar
arXiv AI
Jul 21

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

arXiv:2607. 17017v1 Announce Type: cross Abstract: As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories.

By Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, Yu Shi, Zhongnan Fang, Xuan Cao, Jing Qian, Rui Li
arXiv AI
Aug 24

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

The paper proposes a single hierarchical Semantic ID (SID) system to unify product identification across multiple merchants in e-commerce. By learning SID representations from product content, the authors demonstrate that ranking algorithms can aggregate consumer affinity and product performance over SID prefixes, improving offline relevance and online engagement. For query reformulation, SID concepts guide navigation and refinement, yielding better intent preservation and higher-quality suggestions compared to taxonomy or raw query transitions.

By Steven Xu, Sanjyot Thete, Saathvik Dirisala, Raghav Saboo, Nimesh Sinha, Leo Shao, Elyse Winer, Sudeep Das, Martin Wang, Kyle MacDonald
Hugging Face Trending Papers
Jul 29

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items.

arXiv Machine Learning
1d ago

RPTune: Learned Context Curation for LLM Catalog Search

RPTune is an end‑to‑end framework that improves in‑context catalog search for small merchant businesses by learning to curate product catalogs and fine‑tuning large language models (LLMs) with catalog‑grounded supervision. It uses an encoder‑reorganizer curator to order and prune products based on LLM feedback, and then applies context‑relative rewards during LLM post‑training. Across seven real merchants and 100 complex conversational queries per merchant, RPTune boosts search accuracy by up to 31.4 percentage points from curation alone and an additional 10.3 points on average from post‑training.

By Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan \"O. Ar{\i}k