ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
arXiv:2512. 16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community.
SynIL is a new framework for offline imitation learning that automatically assesses the quality of demonstration data without requiring labels. It uses motor synergy—a low‑dimensional coordinated movement pattern linked to proficiency—to generate dense, transition‑level reward signals through self‑supervised reward regression. Experiments on D4RL locomotion and Robomimic manipulation datasets show that synergy‑derived rewards align well with true rewards and that SynIL outperforms Behavior Cloning and rivals or surpasses offline reinforcement learning in sparse‑reward scenarios.
arXiv:2512. 16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community.
DexPIE is a post‑training framework that improves dexterous manipulation policies using real‑world experience. It introduces a dexterous‑hand‑adapted intervention system and multi‑stage DAgger‑style data collection to enhance exploration, aligns training and inference to reduce distribution shift, and conditions the policy on a continuous optimality indicator for fine‑grained data quality use. In three real‑world tasks, DexPIE boosts success rates by 37.3% over a demonstration‑based baseline, outperforming all other methods and showing stronger robustness.
arXiv:2505. 04999v2 Announce Type: replace-cross Abstract: Learning robot control policies from demonstrations typically requires action-labeled expert data, which is expensive to collect through teleoperation.
arXiv:2608. 11363v1 Announce Type: cross Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction.
arXiv:2606. 09758v1 Announce Type: cross Abstract: Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment.
arXiv:2607. 02466v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale.
arXiv:2607. 01225v1 Announce Type: cross Abstract: Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights.
arXiv:2607. 15880v1 Announce Type: cross Abstract: Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization.
arXiv:2607. 07859v1 Announce Type: new Abstract: Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values.
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
arXiv:2609.37599v1 Announce Type: cross Abstract: Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators...
arXiv:2502. 19544v3 Announce Type: replace Abstract: Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL).