arXiv Machine Learning

Variance Reduction for Heavy-Tailed Monetization Metrics in Ranking Experiments via Post-Stratification

arXiv:2606. 04110v1 Announce Type: new Abstract: Online evaluation of ranking and retrieval systems often relies on downstream monetization metrics such as app revenue or creator earnings.

arXiv AI
6d ago

SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations

SPADE (Serendipitous Pareto Distance Evaluation) is a new metric for recommender systems that simultaneously considers item similarity, popularity, and user relevance. It projects items into a two‑dimensional space and computes a user‑specific Pareto frontier of maximally popular and historically similar items, then averages the minimum Euclidean distance from this frontier for correctly recommended test‑set items. Experiments on five datasets and five baseline algorithms demonstrate that SPADE effectively discourages algorithms from exploiting accuracy‑only metrics and reliably isolates serendipitous discoveries.

By Tobias Vente, Maarten Peirsman, Noah Dani\"els, Hannu Toivonen, Bart Goethals
Hugging Face Trending Papers
Aug 27

Incremental Recommendation via Causal Models

The paper proposes an incremental recommendation approach that uses a causal model built from existing holdback data to avoid delivering redundant recommendations. By applying a dual‑threshold targeting policy, the system only recommends content when the likelihood of a treated stream is high and the likelihood of an organic stream is low, thereby reducing recommendation impressions by 7% without hurting overall consumption. Joint training with holdback data also improves the calibration of the treated head, suggesting that causal models capture more generalisable representations than purely observational models.

arXiv Computation and Language
6d ago

ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker

ZooWork-ShopRanker is a family of open e‑commerce rerankers (0.6B, 4B, and 8B) that align with human shopping preferences by using large language models as preference oracles to generate training pairs. The flagship 8B model serves as a teacher for the smaller 4B and 0.6B models, which are further refined on judged pairs. A new benchmark, ShopRank‑Bench, contains ~10,000 private‑traffic preference pairs and shows that all ZooWork models outperform the strongest open reranker baseline and their own un‑aligned versions.

By Siqiao Xue, Shuxuan Liu, Ning Hu
arXiv Machine Learning
Sep 22

Connected Content Retriever: Dense Graph Edge Features Powering Pre-Ranking at LinkedIn

The paper introduces Connected Content Retriever (CC Retriever), a pre‑ranking system for LinkedIn’s Feed that uses dense graph edge features to score candidate content from a billion‑scale index within a 120 ms latency budget. By leveraging GPU‑based sorted‑search primitives, the system can apply a full deep ranking model with 50× more parameters, achieving a 2.5% lift in content time spent in online experiments. The work details the economic‑graph features and model architecture that enable this scalable, low‑latency scoring pipeline.

By Akhilesh Gupta, Sudarshan Srinivasa Ramanujam, Chirag Bhanuprasad Mehta, Reshma Asharaf Beena, Dhritiman Das, Birjodh Singh Tiwana, Bhargavkumar Kanubhai Patel, Mack Lee, Renyi Tang