← Back to all news
arXiv Machine Learning September 30, 2026 By Gal Mendelson, Eyal Tadmor

Fooling Algorithms in Non-Stationary Bandits using Belief Inertia

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 29

Learning in Markovian bandits with non-observable states and constrained decision epochs

arXiv:2606. 27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs.

By Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop
reinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Aug 18

Toward Optimal Second-Order Path-Length Guarantee for Adversarial Multi-Armed Bandits

arXiv:2608. 15996v1 Announce Type: new Abstract: We study second-order path-length regret in adversarial $K$-armed bandits against oblivious loss sequences.

By Mengxiao Zhang
reinforcement-learningsafety
More like this →
arXiv Machine Learning
Jun 9

Algorithm for Contextual Queueing Bandits with Rate-Optimal Queue Length Regret

arXiv:2606. 09668v1 Announce Type: new Abstract: Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates.

By Seoungbin Bae, Dabeen Lee
More like this →
arXiv Machine Learning
Aug 18

Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits

arXiv:2510. 22819v3 Announce Type: replace Abstract: The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learner's actual decisions and describes the evolution of the learning process over time.

By Jingxin Zhan, Yuze Han, Zhihua Zhang
safety
More like this →
arXiv Machine Learning
Sep 15

Toward Optimal Switching Regret for Multi-Armed Bandits with Oblivious Adversary

arXiv:2609.13547v1 Announce Type: new Abstract: We study switching regret in adversarial multi-armed bandits, where the learner competes with an arm sequence that changes at most $S$ times. When $S$...

By Mengxiao Zhang
safety
More like this →
arXiv Machine Learning
Aug 19

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

arXiv:2608. 17841v1 Announce Type: cross Abstract: Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs.

By Kaifei Wang, Yinyu Ye, Han Zhong
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea