arXiv Machine Learning

A Perturbation Approach to Unconstrained Linear Bandits

arXiv:2603. 28201v3 Announce Type: replace Abstract: We revisit the standard perturbation-based approach of Abernethy et al.

arXiv Machine Learning
Jul 7

Dynamic Regret for Non-Stationary Linear Bandits via Misspecification Reductions

arXiv:2607. 02891v1 Announce Type: new Abstract: Many online decision-making problems involve both round-specific feasible actions and drifting reward models: eligible ad impressions, feasible prices, and available treatments can change over time, while user preferences, demand curves, and patient responses may evolve.

By Zihao Hu, Yuan Yao, Jiheng Zhang, Zhengyuan Zhou
arXiv Machine Learning
Jun 19

Stabilizing Bandits using Regularization: Precise Regret and A Quantitative Central Limit Theorem

arXiv:2603. 10184v2 Announce Type: replace-cross Abstract: Statistical inference with bandit data presents fundamental challenges owing to adaptive sampling, which violates the independence assumptions underlying classical asymptotic theory.

By Budhaditya Halder, Ishan Sengupta, Koustav Chowdhury, Samya Praharaj, Koulik Khamaru
arXiv Machine Learning
Jun 29

Self-Concordant Perturbations for Linear Bandits

arXiv:2510. 24187v3 Announce Type: replace-cross Abstract: We consider the adversarial linear bandits setting and present a unified algorithmic framework that bridges Follow-the-Regularized-Leader (FTRL) and Follow-the-Perturbed-Leader (FTPL) methods, extending the known connection between them from the full-information setting.

By Lucas L\'evy, Jean-Lou Valeau, Arya Akhavan, Patrick Rebeschini
arXiv Machine Learning
Sep 10

Improved Dimension Dependence for Bandit Convex Optimization with Gradient Variations

The paper presents an improved analysis of non‑consecutive gradient variation in Bandit Convex Optimization (BCO) with two‑point feedback, leading to better dimension dependence for both convex and strongly convex functions compared to prior work. It also derives new problem‑dependent guarantees such as gradient‑variance and small‑loss regret bounds, extends the technique to one‑point bandit linear optimization over hyper‑rectangular domains, and establishes the first gradient‑variation dynamic and universal regret bounds for two‑point BCO.

By Hang Yu, Yu-Hu Yan, Peng Zhao
arXiv Machine Learning
Aug 18

Generalized Linear Bandits with Memory

arXiv:2608. 15848v1 Announce Type: cross Abstract: We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix.

By Heesang Ann, Hyunjun Choi, Taehyun Hwang, Younghoon Shin, Haeju Cheong, Min-hwan Oh
arXiv Machine Learning
Jul 9

Nonlinear Bandit

arXiv:2607. 07304v1 Announce Type: new Abstract: In this paper we first study the problem of generalized linear bandit (GLB) under heavy-tailed noise.

By Tianshuo Zheng, Ting Wu, Zhi-Hua Zhou, Keqin Liu