arXiv Machine Learning

Contextual Scalarisation Thompson Sampling for multi-objective decisions in public media

The paper introduces Contextual Scalarisation Thompson Sampler (CSTS), a multi‑objective contextual bandit algorithm that learns to weight competing objectives based on observed context. It addresses the need for adaptable decision‑making in public media, where goals such as audience reach, cultural values, and operational constraints must be balanced. Experiments on Radio Télévision Suisse data demonstrate that CSTS improves contextual relevance and aligns more closely with expert curation than fixed‑weight or standard bandit methods.

arXiv Machine Learning
Aug 18

Sequential Batch Learning in Finite-Action Linear Contextual Bandits

arXiv:2004. 06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end.

By Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou
arXiv AI
Aug 10

Progressive Content Refinement with Decaying Reward Joint LinUCB

arXiv:2608. 06750v1 Announce Type: cross Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect.

By Shion Ishikawa, Pablo Loyola, Young-joo Chung, Yun Ching Liu
arXiv Machine Learning
Aug 13

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan