arXiv:2607. 01225v1 Announce Type: cross Abstract: Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights.
By Chih-Han Yang, Dai-Jie Wu, Yun-Ping Huang, Ping-Chun Hsieh, Kenneth Marino, Shao-Hua Sun
SynIL is a new framework for offline imitation learning that automatically assesses the quality of demonstration data without requiring labels. It uses motor synergy—a low‑dimensional coordinated movement pattern linked to proficiency—to generate dense, transition‑level reward signals through self‑supervised reward regression. Experiments on D4RL locomotion and Robomimic manipulation datasets show that synergy‑derived rewards align well with true rewards and that SynIL outperforms Behavior Cloning and rivals or surpasses offline reinforcement learning in sparse‑reward scenarios.
By Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe
arXiv:2506. 06793v2 Announce Type: replace-cross Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning.
By Zixuan Dong, Yumi Omori, Keith Ross
arXiv:2607. 10601v1 Announce Type: new Abstract: Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitation.
By Yixiong Chen, Alan Yuille
arXiv:2606. 09758v1 Announce Type: cross Abstract: Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment.
By Quinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal, Siddhartha Srinivasa, Abhishek Gupta
SAIL is a framework that transforms robot imitation learning into an iterative refinement problem, enabling test-time scaling of trajectory generation. It employs Monte Carlo Tree Search where each node represents a full trajectory and edges denote refinements, guided by an archive of successful trajectories, a vision‑language model for scoring, and step‑level feedback. Experiments on six manipulation tasks in simulation and real‑world settings show that higher test‑time compute consistently raises success rates, reaching up to 95% on complex tasks.
By Makoto Sato, Yusuke Iwasawa, Yujin Tang, So Kuroki
The paper introduces Q-Target Pretrained Transformers (QTPT), a method that replaces supervised behavior cloning with a Bellman-style Q‑target objective for in‑context reinforcement learning. QTPT retains the context‑conditioned Transformer architecture but learns to estimate action values using rewards and transitions from the context, rather than merely imitating offline actions. The authors provide theoretical analysis in stochastic linear bandits and finite‑horizon MDPs, demonstrating improved robustness to weak or suboptimal data, and empirically show gains over supervised pretraining on controlled RL benchmarks and extensions to D4RL Kitchen and AntMaze.
By Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao
arXiv:2606. 01238v1 Announce Type: cross Abstract: While diffusion-based policies have impressive performance and expressivity, their long offline training slows down the data collection and policy deployment loop.
By Raghav Mishra, Ian R. Manchester
arXiv:2610.01652v1 Announce Type: cross
Abstract: Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but exist...
By Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons
arXiv:2601. 22823v2 Announce Type: replace-cross Abstract: We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions.
By Mathieu Petitbois, R\'emy Portelas, Sylvain Lamprier
arXiv:2607. 21302v1 Announce Type: new Abstract: Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by leveraging policy priors derived from offline demonstrations.
By Gong Gao, Weidong Zhao, Xianhui Liu, Ning Jia
arXiv:2512. 16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community.
By Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett