arXiv:2606. 19752v1 Announce Type: cross Abstract: Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while rare efficient behaviors may be forgotten during training.
By Yinsen Jia, Boyuan Chen
arXiv:2603. 11395v3 Announce Type: replace-cross Abstract: Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving performance in both past and future tasks.
By Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang, Levin Kuhlmann, Gideon Kowadlo
arXiv:2606. 03598v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation.
By Ziyang Chen, Shaoguang Wang, Weiyu Guo, Qianyi Cai, He Zhang, Pengteng Li, Yiren Zhao, Yandong Guo
The paper tackles sample efficiency in image-based reinforcement learning by combining novelty and surprise signals to prioritize experiences. It proposes Novelty and Surprise Prioritized Experience Replay (NSPER) and an extended version, NSPER+R, which also uses these signals as intrinsic rewards. Experiments on DeepMind Control Suite tasks demonstrate that both methods accelerate training and improve convergence compared to existing techniques.
The paper proposes Novelty and Surprise Prioritized Experience Replay (NSPER) for image-based reinforcement learning, combining novelty to highlight underrepresented states and surprise to reveal gaps in the agent’s knowledge. An extended version, NSPER+R, also uses these signals as intrinsic rewards to enhance both replay quality and exploration. Experiments on DeepMind Control Suite tasks demonstrate that NSPER and NSPER+R accelerate training and improve convergence compared to existing methods.
By Hoda Yamani, Henry Williams, Bruce A. MacDonald
arXiv:2512. 16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community.
By Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett
arXiv:2603. 04910v2 Announce Type: replace-cross Abstract: Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition on single-step observations or short-context histories, making them struggle with non-Markovian tasks that require long-term memory.
By Yuheng Lei, Zhixuan Liang, Hongyuan Zhang, Ping Luo
arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.
By Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
EgoSpeedUp is a framework that transfers human manipulation tempo to robot policies by aligning and retiming robot demonstrations using phase-wise tempo estimates derived from human demonstrations. The method improves task success rates by an average of 25 percentage points and reduces successful execution time by 36.5% on two real-world manipulation tasks. It demonstrates that human manipulation tempo can serve as an effective temporal reference for faster and more reliable robot policies.
By Hanbit Oh, Yukiyasu Domae, Takuma Yagi
arXiv:2604. 18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adaptation based on the outcomes of multiple past trials.
By Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song
arXiv:2605. 12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for downstream exploration.
By Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta
arXiv:2607. 05541v1 Announce Type: cross Abstract: Reinforcement Learning is commonly used to train large language models using environmental feedback.
By Muhammad Zain Amin, Kibele Sebnem Yildirim