arXiv Machine Learning

On Incentivized Exploration beyond Bayesianism and Full-Information

arXiv:2607. 18300v1 Announce Type: cross Abstract: We extend Incentive Compatible Exploration beyond the Bayesian full-information setting of Kremer et al.

arXiv Machine Learning
23h ago

Latent Order Bandits

arXiv:2605. 07304v2 Announce Type: replace Abstract: Bandit algorithms solve diverse sequential decision-making problems, but are often too sample-inefficient for from-scratch personalization.

By Emil Carlsson, Newton Mwai, Fredrik D. Johansson
arXiv Machine Learning
6d ago

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan