arXiv Machine Learning By Stephen Pasteris, Rahul Savani, Theodore Turocy

Tracking the Best Strategy in an Extensive-Form Game

Read the original on arXiv Machine Learning →

arXiv:2608. 09501v1 Announce Type: new Abstract: We consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 21

From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

The paper introduces a straightforward framework that transforms dynamic regret minimization into switching regret minimization by constructing an unbiased random sequence for any comparator sequence. Using this reduction, the authors derive dynamic regret bounds for strongly convex and exp-concave losses of “~O(T^{1/3}P_T^{2/3})” and for general convex losses of “O(√{T(1+P_T)})”, matching known minimax optimal results. The approach leverages off-the-shelf switching regret algorithms and controlled variance to achieve these bounds.

By Yibo Wang, Wenhao Yang, Sifan Yang, Yuanyu Wan, Lijun Zhang
arXiv Machine Learning
Jul 9

Nonlinear Bandit

arXiv:2607. 07304v1 Announce Type: new Abstract: In this paper we first study the problem of generalized linear bandit (GLB) under heavy-tailed noise.

By Tianshuo Zheng, Ting Wu, Zhi-Hua Zhou, Keqin Liu