arXiv Machine Learning By Nikita Sevriukov, Anna Barabanova, Uliana Gagarina, Karina Ivanova, Sofiia Kasaeva, Ilya Levin, Marina Sheshukova

Efficient Hypergradient Descent for Inverse Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 3

Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

arXiv:2607. 28849v1 Announce Type: cross Abstract: Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feedback (RL-HF).

By Naman Saxena, Mudit Gaur, Vaneet Aggarwal
arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv AI
Jun 11

Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning

arXiv:2603. 14867v4 Announce Type: replace-cross Abstract: Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcement learning (RL), where a leader agent optimizes its objective while a follower solves a Markov decision process (MDP) conditioned on the leader's decisions.

By Mikoto Kudo, Takumi Tanabe, Akifumi Wachi, Youhei Akimoto
arXiv Machine Learning
4d ago

Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning

The paper introduces Dually Regularized AIL, a model‑free algorithm for adversarial imitation learning that jointly applies KL policy regularization and a quadratic reward penalty based on expert and learner occupancies. It proves fast convergence rates, achieving a ×O(1/K+1/N) bound on the regularized imitation gap in finite‑horizon MDPs with general function approximation, and establishes the first algorithm to attain ×O(1/ε) sample complexity in both expert demonstrations and online interactions for this regularized objective.

By Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang