arXiv AI By Donghwan Lee

Sign-Separated Asymmetric Finite-Time Error Analysis of Q-Learning

Read the original on arXiv AI →

arXiv:2605. 16103v2 Announce Type: replace Abstract: Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, positive errors can be selected and propagated, causing learned values to exceed the true optimal values.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.