arXiv:2609.36421v1 Announce Type: cross
Abstract: Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. T...
By Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake
arXiv:2406.03678v2 Announce Type: replace
Abstract: On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensi...
By Yaozhong Gan, Renye Yan, Zhe Wu, Junliang Xing
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
The paper introduces a reinforcement learning framework that uses groupoids to model local, state-dependent symmetries, allowing agents to discover equivalence structures during interaction. By maintaining orbit representatives and transporters that map raw states to canonical forms, learning and decision-making occur in a symmetry-reduced space while preserving local distinctions. Experiments show that this approach improves sample efficiency and convergence in dense, large-scale environments with partial symmetries, outperforming standard Q‑learning.
By Ben Opperman, Eduardo Alonso, Esther Mondrag\'on
arXiv:2608. 07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly.
By Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, Clare Lyle
Squint is a visual Soft Actor Critic algorithm designed to accelerate reinforcement learning for robotics. It combines parallel simulation, a distributional critic, resolution squinting, layer normalization, a tuned update-to-data ratio, and an optimized implementation to reduce wall‑clock training time. On the SO‑101 Task Set, Squint trains policies in as little as 15 minutes on a single RTX 3090 GPU, with most tasks converging in under 6 minutes and successfully transferring to a real SO‑101 robot.
By Abdulaziz Almuzairee, Henrik I. Christensen