arXiv AI

Language-Critique Imitation Learning from Suboptimal Demonstrations

arXiv:2607. 01225v1 Announce Type: cross Abstract: Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights.

arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv Machine Learning
Jul 8

Supervised Reward Inference

arXiv:2502. 18447v2 Announce Type: replace Abstract: Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models.

By Will Schwarzer, Jordan Schneider, Philip S. Thomas, Scott Niekum
arXiv AI
3d ago

SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

SynIL is a new framework for offline imitation learning that automatically assesses the quality of demonstration data without requiring labels. It uses motor synergy—a low‑dimensional coordinated movement pattern linked to proficiency—to generate dense, transition‑level reward signals through self‑supervised reward regression. Experiments on D4RL locomotion and Robomimic manipulation datasets show that synergy‑derived rewards align well with true rewards and that SynIL outperforms Behavior Cloning and rivals or surpasses offline reinforcement learning in sparse‑reward scenarios.

By Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe