arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
arXiv:2607. 28849v1 Announce Type: cross Abstract: Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feedback (RL-HF).
By Naman Saxena, Mudit Gaur, Vaneet Aggarwal
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv:2609.38955v1 Announce Type: cross
Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a...
By Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu
arXiv:2603. 14867v4 Announce Type: replace-cross Abstract: Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcement learning (RL), where a leader agent optimizes its objective while a follower solves a Markov decision process (MDP) conditioned on the leader's decisions.
By Mikoto Kudo, Takumi Tanabe, Akifumi Wachi, Youhei Akimoto
The paper introduces Dually Regularized AIL, a model‑free algorithm for adversarial imitation learning that jointly applies KL policy regularization and a quadratic reward penalty based on expert and learner occupancies. It proves fast convergence rates, achieving a ×O(1/K+1/N) bound on the regularized imitation gap in finite‑horizon MDPs with general function approximation, and establishes the first algorithm to attain ×O(1/ε) sample complexity in both expert demonstrations and online interactions for this regularized objective.
By Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang
The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.
By Chenyang Cao, Miguel Rogel-Garc\'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart
arXiv:2605. 14599v2 Announce Type: replace-cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) in finite-horizon MDPs with Borel state and action spaces.
By Andreas Schlaginhaufen, Maryam Kamgarpour
The paper introduces a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves the non-uniqueness of reward functions by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. It models expert demonstrations as a Markov chain, estimates the expert policy via penalized maximum likelihood, and provides high-probability bounds on the excess Kullback–Leibler divergence between the estimated and true policies. These results yield non-asymptotic minimax optimal convergence rates for the least-squares reward function, highlighting the trade-offs among smoothing, model complexity, and sample size.
By Denis Belomestny, Alexey Naumov, Artemy Rubtsov, Sergey Samsonov
arXiv:2605. 14982v2 Announce Type: replace-cross Abstract: We address the discounted reward setting in reinforcement learning (RL).
By Sanjeev Manivannan, Shuban V
arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.
By Soichiro Nishimori, Paavo Parmas
arXiv:2607. 08647v1 Announce Type: cross Abstract: As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment.
By Ali Larian, Qian Lin, Chang Zong Wu, Daniel S. Brown