arXiv Machine Learning By Neeti Pokharna, Olivier Jeunen, Yatharth Saraf, Aleksei Ustimenko

Variance Reduction for Heavy-Tailed Monetization Metrics in Ranking Experiments via Post-Stratification

Read the original on arXiv Machine Learning →

arXiv:2606. 04110v1 Announce Type: new Abstract: Online evaluation of ranking and retrieval systems often relies on downstream monetization metrics such as app revenue or creator earnings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations

SPADE (Serendipitous Pareto Distance Evaluation) is a new metric for recommender systems that simultaneously considers item similarity, popularity, and user relevance. It projects items into a two‑dimensional space and computes a user‑specific Pareto frontier of maximally popular and historically similar items, then averages the minimum Euclidean distance from this frontier for correctly recommended test‑set items. Experiments on five datasets and five baseline algorithms demonstrate that SPADE effectively discourages algorithms from exploiting accuracy‑only metrics and reliably isolates serendipitous discoveries.

By Tobias Vente, Maarten Peirsman, Noah Dani\"els, Hannu Toivonen, Bart Goethals
Hugging Face Trending Papers
Aug 27

Incremental Recommendation via Causal Models

The paper proposes an incremental recommendation approach that uses a causal model built from existing holdback data to avoid delivering redundant recommendations. By applying a dual‑threshold targeting policy, the system only recommends content when the likelihood of a treated stream is high and the likelihood of an organic stream is low, thereby reducing recommendation impressions by 7% without hurting overall consumption. Joint training with holdback data also improves the calibration of the treated head, suggesting that causal models capture more generalisable representations than purely observational models.

arXiv Computation and Language
6d ago

ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker

ZooWork-ShopRanker is a family of open e‑commerce rerankers (0.6B, 4B, and 8B) that align with human shopping preferences by using large language models as preference oracles to generate training pairs. The flagship 8B model serves as a teacher for the smaller 4B and 0.6B models, which are further refined on judged pairs. A new benchmark, ShopRank‑Bench, contains ~10,000 private‑traffic preference pairs and shows that all ZooWork models outperform the strongest open reranker baseline and their own un‑aligned versions.

By Siqiao Xue, Shuxuan Liu, Ning Hu