arXiv Machine Learning

Randomized Exploration for Linear Bandits via Absolute Perturbations

arXiv:2606. 28616v1 Announce Type: new Abstract: In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson Sampling (TS) is computationally attractive yet typically harder to analyze due to its non-optimistic nature.

arXiv AI
5d ago

On the Complexity of Preference-Based Bandits

The paper investigates preference-based bandits where a learner selects pairs of arms and receives binary preference feedback modeled by Bradley–Terry. It introduces the locally sensitive eluder dimension, a new complexity measure for logistic preference feedback, and proposes the GINOP algorithm that uses log-loss confidence sets to balance optimism and exploration. The authors prove a first-order regret bound showing that learning with preference feedback can be as statistically efficient as learning from direct rewards, and they validate their theory with empirical experiments.

By Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY)
arXiv Statistics ML
4d ago

Block Optimism for Nonstationary Bandits with Latent Linear Dynamics

The paper introduces a new algorithm for a nonstationary bandit setting where actions influence both immediate rewards and the evolution of an unobserved latent linear state. By approximating the infinite‑memory reward process with a finite‑memory block‑level proxy and applying a UCB‑based block algorithm, the authors achieve a regret bound of “~O(√T)”, improving upon the previous “~O(T^{2/3})” guarantee. This represents the first such “~O(√T)” result for latent linear‑dynamics bandits with bilinear rewards and an open‑loop action‑sequence benchmark.

By Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh
arXiv Machine Learning
Jun 5

Exploration via linearly perturbed loss minimisation

arXiv:2311. 07565v3 Announce Type: replace Abstract: We introduce exploration via linear loss perturbations (EVILL), a randomised exploration method for structured stochastic bandit problems that works by solving for the minimiser of a linearly perturbed regularised negative log-likelihood function.

By David Janz, Shuai Liu, Alex Ayoub, Csaba Szepesv\'ari