arXiv Machine Learning

Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models

arXiv Machine Learning
Sep 11

Statistical analysis of Inverse Entropy-regularized Reinforcement Learning

The paper introduces a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves the non-uniqueness of reward functions by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. It models expert demonstrations as a Markov chain, estimates the expert policy via penalized maximum likelihood, and provides high-probability bounds on the excess Kullback–Leibler divergence between the estimated and true policies. These results yield non-asymptotic minimax optimal convergence rates for the least-squares reward function, highlighting the trade-offs among smoothing, model complexity, and sample size.

By Denis Belomestny, Alexey Naumov, Artemy Rubtsov, Sergey Samsonov
arXiv AI
1d ago

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning

The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.

By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv Machine Learning
Jul 17

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

arXiv:2607. 14373v1 Announce Type: new Abstract: We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics.

By Yang Liu, Yuhao Liu, Yunran Wei
arXiv Machine Learning
Jul 17

A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

arXiv:2607. 14522v1 Announce Type: new Abstract: We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC).

By Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang