← Back to all news
arXiv Machine Learning September 15, 2026 By Tianyuan Jin

Nearly Minimax-Optimal Regret for Linear Contextual Bandits with Arbitrary Adaptive Action Sets

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 18

Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits

arXiv:2608. 15996v1 Announce Type: new Abstract: We study second-order path-length regret in adversarial $K$-armed bandits against oblivious loss sequences.

By Mengxiao Zhang
reinforcement-learningsafety
More like this →
arXiv Machine Learning
Aug 27

Minimax Alternating Regret for the Experts Problem and Online Convex Optimization

arXiv:2608. 25182v1 Announce Type: cross Abstract: In this paper, we study alternating regret in online convex optimization (OCO), motivated by the success of alternating learning dynamics in two-player games.

By Mengxiao Zhang
More like this →
arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
arXiv Machine Learning
Aug 26

Optimal Alternating Regret for Online Learning and Games

arXiv:2608.24731v1 Announce Type: new Abstract: We settle the minimax-optimal alternating regret, a regret notion motivated by alternating learning dynamics in games, for both online linear optimizat...

By Yixin Tao, Weiqiang Zheng
More like this →
arXiv Machine Learning
Aug 21

The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex Optimization

arXiv:2607. 18652v3 Announce Type: replace-cross Abstract: We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for $1$-Lipschitz functions on the $d$-dimensional Euclidean ball.

By Nived Rajaraman, Yanjun Han
llmsreinforcement-learning
More like this →
arXiv Machine Learning
Jul 23

Breaking the $T^{3/4}$ Barrier for Regret Minimization With Bi-Dimensional CDFs

arXiv:2607. 20258v1 Announce Type: new Abstract: We study regret minimization for learning CDF-related objectives of the form \[ g(x)\cdot\mathbb{P}_{X\sim\mathcal{D}}(X\le x), \] over $[0,1]^2$, where $g$ is a known Lipschitz function and $\mathcal{D}$ is an unknown distribution.

By Matteo Castiglioni, Anna Lunghi, Alberto Marchesi
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea