arXiv Machine Learning

Contextual Bandits for Maximizing Stimulated Word-of-Mouth Rewards

arXiv:2606. 15146v1 Announce Type: new Abstract: Stimulated word-of-mouth is a strategy that promotes information sharing through prompts or incentives.

arXiv AI
4d ago

Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

arXiv:2609.37800v1 Announce Type: cross Abstract: Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two cha...

By Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ek\c{s}i, Rahul Sharma, Julia Mueller, Theresa Dombrowski, Jakob Karolus, Viktor Bengs, Eyke H\"ullermeier, Sebastian Vollmer
arXiv Machine Learning
Sep 15

Contextual Scalarisation Thompson Sampling for multi-objective decisions in public media

The paper introduces Contextual Scalarisation Thompson Sampler (CSTS), a multi‑objective contextual bandit algorithm that learns to weight competing objectives based on observed context. It addresses the need for adaptable decision‑making in public media, where goals such as audience reach, cultural values, and operational constraints must be balanced. Experiments on Radio Télévision Suisse data demonstrate that CSTS improves contextual relevance and aligns more closely with expert curation than fixed‑weight or standard bandit methods.

By Th\'eo Ma\"etz, Luc Guillet, Andrea Cavallaro
arXiv Machine Learning
Jul 17

Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning

arXiv:2607. 14192v1 Announce Type: new Abstract: As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention.

By Dingsu Wang, Filip Ryzner, Kelly He, Armando Ordorica, David Woo, Aditya Mantha, Liyao Lu, Usha Amrutha Nookala, Haoran Guo, Jiacong He, Olafur Gudmundsson, Matt Chun, Krystal Benitez, Dhruvil Deven Badani, Yijie Dylan Wang
arXiv AI
Aug 10

Progressive Content Refinement with Decaying Reward Joint LinUCB

arXiv:2608. 06750v1 Announce Type: cross Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect.

By Shion Ishikawa, Pablo Loyola, Young-joo Chung, Yun Ching Liu