arXiv AI

ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning

arXiv:2512. 16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community.

arXiv AI
Aug 19

ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

The paper introduces ORPA, a framework that adds a lightweight, feedback-conditioned module to a pretrained robotic manipulation policy, enabling real‑time residual adjustments in joint space without retraining the base policy. ORPA allows immediate correction of execution errors and distribution shifts, improving success rates and recovery on precision‑sensitive tasks compared to baseline policies and rule‑based inverse kinematics. The method is evaluated on the ALOHA platform, showing its effectiveness in real‑time deployment scenarios.

By Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae
arXiv AI
Jun 2

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

arXiv:2603. 15956v3 Announce Type: replace-cross Abstract: Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data.

By Zifan Xu, Ran Gong, Maria Vittoria Minniti, Kausik Sivakumar, Ahmet Salih Gundogdu, Eric Rosen, Riedana Yan, Tushar Kusnur, Zixing Wang, Di Deng, Peter Stone, Xiaohan Zhang, Karl Schmeckpeper
arXiv AI
3d ago

SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

SynIL is a new framework for offline imitation learning that automatically assesses the quality of demonstration data without requiring labels. It uses motor synergy—a low‑dimensional coordinated movement pattern linked to proficiency—to generate dense, transition‑level reward signals through self‑supervised reward regression. Experiments on D4RL locomotion and Robomimic manipulation datasets show that synergy‑derived rewards align well with true rewards and that SynIL outperforms Behavior Cloning and rivals or surpasses offline reinforcement learning in sparse‑reward scenarios.

By Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe
arXiv Computer Vision
Sep 21

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

DexPIE is a post‑training framework that improves dexterous manipulation policies using real‑world experience. It introduces a dexterous‑hand‑adapted intervention system and multi‑stage DAgger‑style data collection to enhance exploration, aligns training and inference to reduce distribution shift, and conditions the policy on a continuous optimality indicator for fine‑grained data quality use. In three real‑world tasks, DexPIE boosts success rates by 37.3% over a demonstration‑based baseline, outperforming all other methods and showing stronger robustness.

By Ruizhe Liao, Wenrui Chen, Liangji Zeng, Haoran Lin, Fan Yang, Kailun Yang, Yaonan Wang
arXiv Machine Learning
Aug 4

DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration

arXiv:2608. 01452v1 Announce Type: cross Abstract: Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that are moving or require rapid adjustments.

By Haoran Liao, Pengyue Wang, Shuoyu Chen, Kehan Cheng, Xuhang Chen, Yuhao Lin, Mu Lin, Zhizhao Liang, Xiaoyi Fan, Chengyi Xing, Dan Niu, Yi-Lin Wei, Wei-Shi Zheng
arXiv Computer Vision
Sep 25

EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies

EgoSpeedUp is a framework that transfers human manipulation tempo to robot policies by aligning and retiming robot demonstrations using phase-wise tempo estimates derived from human demonstrations. The method improves task success rates by an average of 25 percentage points and reduces successful execution time by 36.5% on two real-world manipulation tasks. It demonstrates that human manipulation tempo can serve as an effective temporal reference for faster and more reliable robot policies.

By Hanbit Oh, Yukiyasu Domae, Takuma Yagi
arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv Machine Learning
Sep 21

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

The paper introduces PARTS, a real‑world subtask reinforcement learning framework that fine‑tunes a pretrained robot policy by focusing on critical bottleneck subtasks while keeping the base policy frozen. It uses agent‑generated selectors and success verifiers to provide local rewards, enabling learning even when full‑task successes are rare. Experiments on bimanual YAM and single‑arm Franka robots show that PARTS raises complete‑task success from 32% to 61% and from 50% to 95%, respectively, with only tens of minutes of real‑world RL rollouts and minimal human intervention.

By Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang, Lingfeng Sun
arXiv AI
6d ago

SUN: Agentic Robot Policy Learning with Persistent Task Programs

The paper introduces SUN (Semantically UNified) Programs, typed executables that translate grounded relations into optimal control objectives, satisfaction predicates, and learning rewards. Using the Kuafu harness, a foundation model orchestrates scene preparation, verification, residual reinforcement learning, and data generation, repairing candidate programs and calibrating reward weights. Across nine multi‑stage manipulation tasks, Kuafu achieves an 82.03% success rate, outperforms learned baselines, generates demonstrations 10.57× faster than human teleoperation, and transfers zero‑shot to physical Franka and Kinova robots.

By Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong