Efficient Offline Learning of Ranking Policies via Top-$k$ Policy Decomposition
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.
arXiv:2509. 03456v2 Announce Type: replace-cross Abstract: Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits.
arXiv:2506. 06989v3 Announce Type: replace-cross Abstract: Learning-to-rank (LTR) systems commonly depend on implicit feedback, such as user clicks, because it is easy to collect and can serve as a valuable signal of user preferences.
arXiv:2606. 14929v1 Announce Type: cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.
arXiv:2601. 21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards.
arXiv:2609.13730v1 Announce Type: new Abstract: Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work h...