arXiv Machine Learning By Jikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy, Baoyi Shi, Congshan Zhang

Adaptive Exploration for Latent-State Bandits

Read the original on arXiv Machine Learning →

arXiv:2602. 05139v3 Announce Type: replace Abstract: We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.