arXiv Machine Learning

Personalizing Marketplace Policies with Competing Objectives and Constrained Experiments: Evidence from a Job Marketplace

arXiv:2606. 30932v1 Announce Type: new Abstract: Two-sided marketplaces connect distinct user groups whose interests often conflict -- improving outcomes on one side could degrade the other side's experience.

arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv Machine Learning
Aug 13

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
Hugging Face Trending Papers
Aug 12

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.

arXiv AI
4d ago

Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

arXiv:2609.37800v1 Announce Type: cross Abstract: Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two cha...

By Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ek\c{s}i, Rahul Sharma, Julia Mueller, Theresa Dombrowski, Jakob Karolus, Viktor Bengs, Eyke H\"ullermeier, Sebastian Vollmer
arXiv Computation and Language
Sep 15

Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation

arXiv:2609.14648v1 Announce Type: new Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often...

By Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull, Caitlin A. Stamatis, Olivier Tieleman, Guilherme B. Freire, Jinghong Chen
arXiv AI
Sep 24

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.

By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
arXiv AI
Jul 24

Benchmarking the Personalization Capabilities of Large Language Models

arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.

By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv Machine Learning
Jul 31

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.

By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao