Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
arXiv:2605. 23415v2 Announce Type: replace Abstract: Reinforcement learning has long struggled with poor sample efficiency.
arXiv:2605. 23415v2 Announce Type: replace Abstract: Reinforcement learning has long struggled with poor sample efficiency.
The paper introduces a reinforcement learning framework that uses groupoids to model local, state-dependent symmetries, allowing agents to discover equivalence structures during interaction. By maintaining orbit representatives and transporters that map raw states to canonical forms, learning and decision-making occur in a symmetry-reduced space while preserving local distinctions. Experiments show that this approach improves sample efficiency and convergence in dense, large-scale environments with partial symmetries, outperforming standard Q‑learning.
The paper proposes using category theory to structure reinforcement learning in high‑dimensional, partially observable environments. By partitioning the state space into equivalence classes (symmetry orbits) and treating each class as a groupoid with a canonical representative, the agent can share learning across similar states, reducing redundancy and improving sample efficiency. Experiments on partially observable benchmarks show that this orbit‑based partitioning consistently enhances performance in environments with latent symmetry.
arXiv:2607. 11624v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency.
Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency. In robotics, a recent line of work has emerged addressing this problem by encoding physics priors in the learning process.
arXiv:2505. 03296v2 Announce Type: replace-cross Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy representation and imitation learning in robot manipulation.
Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction. Yet prevalent diffusion- and flow-matching robot policies lack tractable likelihoods, limiting their use in likelihood-based offline RL post-training.
arXiv:2510. 11103v3 Announce Type: replace-cross Abstract: Many robotic control tasks require policies to act on orientations, yet the geometry of SO(3) makes this nontrivial.
arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.
arXiv:2606. 01098v1 Announce Type: cross Abstract: Generative action policies based on diffusion or flow matching excel in behavior cloning, yet their iterative sampling is prohibitive for high-frequency robot control.
arXiv:2608. 20208v1 Announce Type: new Abstract: Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction.
arXiv:2607. 10169v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities.