arXiv Machine Learning By Denis Belomestny, Alexey Naumov, Artemy Rubtsov, Sergey Samsonov

Statistical analysis of Inverse Entropy-regularized Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves the non-uniqueness of reward functions by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. It models expert demonstrations as a Markov chain, estimates the expert policy via penalized maximum likelihood, and provides high-probability bounds on the excess Kullback–Leibler divergence between the estimated and true policies. These results yield non-asymptotic minimax optimal convergence rates for the least-squares reward function, highlighting the trade-offs among smoothing, model complexity, and sample size.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning

The paper introduces Dually Regularized AIL, a model‑free algorithm for adversarial imitation learning that jointly applies KL policy regularization and a quadratic reward penalty based on expert and learner occupancies. It proves fast convergence rates, achieving a ×O(1/K+1/N) bound on the regularized imitation gap in finite‑horizon MDPs with general function approximation, and establishes the first algorithm to attain ×O(1/ε) sample complexity in both expert demonstrations and online interactions for this regularized objective.

By Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang