arXiv Machine Learning By AmirHossein Naghdi, Ali Baheri

Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits

Read the original on arXiv Machine Learning →

arXiv:2606. 23933v1 Announce Type: cross Abstract: We study non-stationary linear contextual bandits where the reward model drifts over time, rendering classical contextual bandit algorithms brittle because historical data becomes systematically biased.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.