arXiv Machine Learning

Stochastic Linear Bandits with Partially Observed Actions

arXiv:2607. 08971v1 Announce Type: new Abstract: The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making.

arXiv AI
Jun 30

Exploration and Online Transfer with Behavioral Foundation Models

arXiv:2606. 29980v1 Announce Type: new Abstract: Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories.

By Louis Bagot (SyCoSMA), Mathieu Lefort (LIRIS, SyCoSMA, IRISA, MALT, UR), La\"etitia Matignon (SyCoSMA)
arXiv AI
3d ago

On the Complexity of Preference-Based Bandits

The paper investigates preference-based bandits where a learner selects pairs of arms and receives binary preference feedback modeled by Bradley–Terry. It introduces the locally sensitive eluder dimension, a new complexity measure for logistic preference feedback, and proposes the GINOP algorithm that uses log-loss confidence sets to balance optimism and exploration. The authors prove a first-order regret bound showing that learning with preference feedback can be as statistically efficient as learning from direct rewards, and they validate their theory with empirical experiments.

By Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY)
arXiv Machine Learning
Aug 28

Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions

The paper introduces a new approach to safety in contextual bandits with continuous actions, focusing on high‑probability constraints on the realized cost rather than expected cost. It presents the High‑Probability Constrained UCB algorithm, which balances reward exploration with conservative safety estimation, and provides theoretical regret guarantees for linear models and extensions to general function classes. Experiments demonstrate that this realized‑cost safety framework significantly reduces safety violations compared to expected‑cost constrained methods.

By Spyros Dragazis, Aldo Pacchiano
arXiv Machine Learning
Aug 18

Sequential Batch Learning in Finite-Action Linear Contextual Bandits

arXiv:2004. 06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end.

By Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou
arXiv Machine Learning
Aug 19

Latent Order Bandits

The paper introduces Latent Order Bandits (LOB), a new bandit framework that relaxes the strict assumptions of traditional latent bandits by only requiring a partial order of action preferences within each latent state. LOB allows instances sharing the same state to have different reward distributions as long as the action ranking remains consistent, making it suitable for scenarios like user groups on streaming services who agree on genre preferences but rate differently. The authors present an upper‑confidence bound algorithm for both total and partial latent orders, provide regret bounds, and propose a posterior‑sampling variant that empirically outperforms full‑prior latent bandits when reward scales vary across instances sharing the same latent state.

By Emil Carlsson, Newton Mwai, Fredrik D. Johansson
arXiv Machine Learning
Sep 16

Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

Meta-LinEXP3 is an online-within-online algorithm designed for adversarial linear contextual bandits with random action sets. It builds a task-level prior from completed tasks to guide an inner LinEXP3 learner, achieving an σO(√n) per‑task regret when context distributions are known and an σO(n^{2/3}) regret with a past‑only regularized moment estimator when they are unknown. The paper also links prior accuracy to transfer regret, showing that better priors yield sublinear, transfer‑dependent regret across tasks, and demonstrates the method on structured hyperspectral tensor sampling.

By Hao Li, Jie Xu, Zheng Xie
Hugging Face Trending Papers
Aug 27

Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions

The paper introduces a new approach to safety in contextual bandits with continuous actions by enforcing high‑probability constraints on the realized cost rather than on its expectation. It proposes the High‑Probability Constrained UCB algorithm, which balances optimistic reward exploration with pessimistic safety estimation. The authors provide theoretical regret guarantees for linear models and extend the analysis to general function classes, demonstrating experimentally that realized‑cost constraints significantly reduce safety violations compared to expected‑cost baselines.