arXiv AI By Yinsen Jia, Boyuan Chen

Temporal Self-Imitation Learning

Read the original on arXiv AI →

arXiv:2606. 19752v1 Announce Type: cross Abstract: Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 25

EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies

EgoSpeedUp is a framework that transfers human manipulation tempo to robot policies by aligning and retiming robot demonstrations using phase-wise tempo estimates derived from human demonstrations. The method improves task success rates by an average of 25 percentage points and reduces successful execution time by 36.5% on two real-world manipulation tasks. It demonstrates that human manipulation tempo can serve as an effective temporal reference for faster and more reliable robot policies.

By Hanbit Oh, Yukiyasu Domae, Takuma Yagi
arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters