Impatient Bandits: Optimizing for the Long-Term Without Delay
arXiv:2501. 07761v2 Announce Type: replace-cross Abstract: Increasingly, recommender systems are tasked with improving users' long-term satisfaction.
arXiv:2606. 09891v1 Announce Type: cross Abstract: Ranking in digital marketplaces is a dynamic exposure-allocation mechanism: displayed items shape discovery trajectories and success events logged by the platform to update future allocation policies.
arXiv:2501. 07761v2 Announce Type: replace-cross Abstract: Increasingly, recommender systems are tasked with improving users' long-term satisfaction.
arXiv:2607. 14161v1 Announce Type: cross Abstract: Pinterest is where people turn inspiration into action as users browse ideas, then take steps toward realization, often by discovering shoppable content.
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).
arXiv:2609.35783v1 Announce Type: cross Abstract: Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such enviro...
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge.
The paper introduces “PACE”, a training‑free framework that tackles bottlenecks in Retrieval‑Augmented Generation by frontloading evidence and adaptively budgeting reranking. It first reorders candidate documents based on marginal evidence coverage—prioritizing query‑relevant, complementary, and chain‑forming documents—providing a $(1-1/e)$ approximation guarantee. Then it dynamically adjusts the reranking budget according to the relative pressure of the reranker and the language model, improving evidence recall and reducing p95 latency in multi‑hop QA workloads.
The paper introduces a retrieval‑grounded credit‑assignment method for generative recommenders that use Semantic IDs (SIDs). By structuring each autoregressive trace into a history summary, a set of interest hypotheses, and a final SID, a frozen retriever verifies each hypothesis as a catalog query. Rewards are assigned at the hypothesis level when any query retrieves the target within the top‑K, allowing distinct updates for rollouts that share the same SID reward and improving SID recommendation performance on Amazon Reviews datasets.
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
FairDiff is a new fairness‑aware diffusion framework designed to mitigate the self‑reinforcing Matthew Effect in Diffusion Recommender Models (DRMs). It introduces Popularity Condition Guidance (PCG) to reweight inference‑time gradients and penalize high‑popularity items, and a Semantic Calibration (SC) module that aligns forward and reverse distributions via optimal transport. Experiments show FairDiff achieves state‑of‑the‑art performance while reducing popularity bias in DRMs.
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.
arXiv:2608.28931v1 Announce Type: cross Abstract: Matching users to interest categories at scale is central to personalized shopping, but the task is challenging in large e-commerce platforms, where...