arXiv:2609.13564v1 Announce Type: new
Abstract: We study KL-regularized contextual bandits under both reward and preference feedback. We show that greedy sampling can achieve logarithmic regret witho...
By Zichen Wang, Haoyang Hong, Huazheng Wang
arXiv:2605. 09454v2 Announce Type: replace-cross Abstract: We study the $\textit{single-index bandit}$ problem, where rewards depend on an unknown one-dimensional projection of high-dimensional contexts through an unknown reward function.
By Devdan Dey, Sujoy Bhore, Avishek Ghosh
arXiv:2606. 28616v1 Announce Type: new Abstract: In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson Sampling (TS) is computationally attractive yet typically harder to analyze due to its non-optimistic nature.
By Toshinori Kitamura, Shuai Liu, Csaba Szepesv\'ari
arXiv:2602. 09456v2 Announce Type: replace Abstract: We propose an algorithmic framework, Offline Estimation to Decisions (OE2D), that efficiently reduces contextual bandit learning with general reward function approximation to offline regression.
By Hao Qin, Chicheng Zhang
arXiv:2607. 08979v1 Announce Type: new Abstract: We study the active learning problem of fixed-confidence top-$k$ identification from noisy pairwise comparisons.
By Motti Goldberger, Nils Rudi
arXiv:2609.37209v1 Announce Type: new
Abstract: Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairw...
By Junghyun Lee, Minsoo Ha, Sanghwa Kim, Yeongjong Kim, Eunjee Lee, Seiyun Shin, Kwang-Sung Jun