arXiv:2609.36421v1 Announce Type: cross
Abstract: Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. T...
By Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake
arXiv:2406.03678v2 Announce Type: replace
Abstract: On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensi...
By Yaozhong Gan, Renye Yan, Zhe Wu, Junliang Xing
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
The paper introduces a reinforcement learning framework that uses groupoids to model local, state-dependent symmetries, allowing agents to discover equivalence structures during interaction. By maintaining orbit representatives and transporters that map raw states to canonical forms, learning and decision-making occur in a symmetry-reduced space while preserving local distinctions. Experiments show that this approach improves sample efficiency and convergence in dense, large-scale environments with partial symmetries, outperforming standard Q‑learning.
By Ben Opperman, Eduardo Alonso, Esther Mondrag\'on
arXiv:2608. 07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly.
By Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, Clare Lyle
Squint is a visual Soft Actor Critic algorithm designed to accelerate reinforcement learning for robotics. It combines parallel simulation, a distributional critic, resolution squinting, layer normalization, a tuned update-to-data ratio, and an optimized implementation to reduce wall‑clock training time. On the SO‑101 Task Set, Squint trains policies in as little as 15 minutes on a single RTX 3090 GPU, with most tasks converging in under 6 minutes and successfully transferring to a real SO‑101 robot.
By Abdulaziz Almuzairee, Henrik I. Christensen
arXiv:2609.34851v2 Announce Type: replace
Abstract: Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprec...
By Nam Hee Kim, Markus Kirjonen, Perttu H\"am\"al\"ainen
arXiv:2510. 11103v3 Announce Type: replace-cross Abstract: Many robotic control tasks require policies to act on orientations, yet the geometry of SO(3) makes this nontrivial.
By Martin Schuck, Sherif Samy, Angela P. Schoellig
arXiv:2605. 03065v2 Announce Type: replace Abstract: Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning.
By Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Dai, Paarth Shah, Max Simchowitz
arXiv:2607. 26985v1 Announce Type: cross Abstract: Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times.
By Gabe Everett, Brice Gunter, Ryan Vander Stelt, Cleiver Ruiz-Martinez, Blake Hull, Juan Rojas
The paper proposes using category theory to structure reinforcement learning in high‑dimensional, partially observable environments. By partitioning the state space into equivalence classes (symmetry orbits) and treating each class as a groupoid with a canonical representative, the agent can share learning across similar states, reducing redundancy and improving sample efficiency. Experiments on partially observable benchmarks show that this orbit‑based partitioning consistently enhances performance in environments with latent symmetry.
By Ben Opperman, Eduardo Alonso, Esther Mondrag\'on
arXiv:2606. 27475v1 Announce Type: cross Abstract: Robots trained on real world data tend to be imprecise, slow, and brittle to perturbations.
By Raymond Yu, William Huey, Mustafa Mukadam, Anusha Nagabandi, Abhishek Gupta