arXiv AI

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

arXiv:2603. 15956v3 Announce Type: replace-cross Abstract: Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data.

arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv AI
Aug 13

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

arXiv:2605. 12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for downstream exploration.

By Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta
arXiv Computer Vision
Sep 21

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

DexPIE is a post‑training framework that improves dexterous manipulation policies using real‑world experience. It introduces a dexterous‑hand‑adapted intervention system and multi‑stage DAgger‑style data collection to enhance exploration, aligns training and inference to reduce distribution shift, and conditions the policy on a continuous optimality indicator for fine‑grained data quality use. In three real‑world tasks, DexPIE boosts success rates by 37.3% over a demonstration‑based baseline, outperforming all other methods and showing stronger robustness.

By Ruizhe Liao, Wenrui Chen, Liangji Zeng, Haoran Lin, Fan Yang, Kailun Yang, Yaonan Wang
arXiv AI
Sep 17

Missing Bridges: Composition-Aware Active Imitation Learning

The paper introduces Adaptive Agents via Latent Topologies (AALT), a method for active imitation learning that selects demonstrations based on their expected impact on start‑to‑goal connectivity rather than generic information gain. AALT builds a latent topology of hub states and learned behaviors, identifies high‑value bridge demonstrations that can solve many tasks simultaneously, and uses these to condition a diffusion policy for planning. In a simulated UR5e robot retrieval task with 72 start‑goal pairs, AALT achieved 100% success after only three demonstrations, outperforming baselines that required many more queries.

By Maxwell J. Jacobson, Ahmed H Qureshi, Yexiang Xue
Hugging Face Trending Papers
Aug 20

RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation

Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction. Yet prevalent diffusion- and flow-matching robot policies lack tractable likelihoods, limiting their use in likelihood-based offline RL post-training.