arXiv AI By Jayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa

ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning

Read the original on arXiv AI →

ADEPT is a reinforcement‑learning framework that first pre‑trains a dexterous policy on a generic object reposing task and then post‑trains downstream policies using this pretrained behavior as a prior. The approach avoids relearning basic skills for each new task, and employs a stable post‑training recipe—behavior‑cloning distillation, critic warm‑up, and conservative on‑policy updates—to preserve the pretrained capabilities. ADEPT’s joint‑space Geometric Fabric mediates between the policy and the robot, enabling zero‑shot sim‑to‑real transfer on a 23‑DoF Kuka‑Allegro and a 29‑DoF Flexiv‑Sharpa, where the robots solve long‑horizon tasks from challenging initial states at human‑level speed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 13

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

arXiv:2605. 12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for downstream exploration.

By Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta