arXiv:2609.15170v1 Announce Type: new
Abstract: We study stochastic linear contextual bandits with arbitrary action menus that may depend on the fixed parameter and the interaction history. We establ...
By Tianyuan Jin
arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.
By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
arXiv:2609.13547v1 Announce Type: new
Abstract: We study switching regret in adversarial multi-armed bandits, where the learner competes with an arm sequence that changes at most $S$ times. When $S$...
By Mengxiao Zhang
arXiv:2609. 17595v1 Announce Type: new Abstract: In the improving multi-armed bandits problem, each of $k$ arms has an unknown nondecreasing, discretely concave reward curve $f_i$, and pulling arm $i$ for the $t$-th time yields $f_i(t)$.
By Xuan Li
arXiv:2609. 38659v1 Announce Type: cross Abstract: We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers.
By Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu
arXiv:2603. 25029v4 Announce Type: replace Abstract: We study online convex optimization (OCO) with two-point bandit feedback against a non-anticipating adaptive adversary.
By Haishan Ye
arXiv:2609.38659v2 Announce Type: replace-cross
Abstract: We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multi...
By Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu
arXiv:2608. 15365v1 Announce Type: new Abstract: Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits.
By Jingxin Zhan, Yuze Han, Zhihua Zhang
arXiv:2510. 22819v3 Announce Type: replace Abstract: The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learner's actual decisions and describes the evolution of the learning process over time.
By Jingxin Zhan, Yuze Han, Zhihua Zhang
arXiv:2511.05620v2 Announce Type: replace
Abstract: We study worst-case dynamic regret of specific multi-armed bandit algorithms on piecewise-stationary instances with at most one breakpoint. Our con...
By Gal Mendelson, Eyal Tadmor
arXiv:2608. 17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs.
By Kaifei Wang, Yinyu Ye, Han Zhong
arXiv:2606. 09668v1 Announce Type: new Abstract: Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates.
By Seoungbin Bae, Dabeen Lee