arXiv AI

Categorical Internalisation of Environmental Groupoids for Generalisable POMDP Solving

The paper proposes using category theory to structure reinforcement learning in high‑dimensional, partially observable environments. By partitioning the state space into equivalence classes (symmetry orbits) and treating each class as a groupoid with a canonical representative, the agent can share learning across similar states, reducing redundancy and improving sample efficiency. Experiments on partially observable benchmarks show that this orbit‑based partitioning consistently enhances performance in environments with latent symmetry.

arXiv Machine Learning
Sep 14

Groupoid-Based Internal State Representations for Reinforcement Learning with Local Symmetries

The paper introduces a reinforcement learning framework that uses groupoids to model local, state-dependent symmetries, allowing agents to discover equivalence structures during interaction. By maintaining orbit representatives and transporters that map raw states to canonical forms, learning and decision-making occur in a symmetry-reduced space while preserving local distinctions. Experiments show that this approach improves sample efficiency and convergence in dense, large-scale environments with partial symmetries, outperforming standard Q‑learning.

By Ben Opperman, Eduardo Alonso, Esther Mondrag\'on
arXiv Machine Learning
Jun 9

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

arXiv:2601. 22211v2 Announce Type: replace Abstract: Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical.

By Lingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti, Andrew Ma, Wenbo Chen, Mingxiao Song, Lily Xu, Milind Tambe
arXiv AI
Aug 25

Reward-Free Continual Adaptation for Resilient Space Robots

The paper presents a reward‑free continual learning framework for space robots that uses latent‑state world models to adapt to severe hardware degradation. By pre‑training a model‑based agent in diverse simulations, the world model learns to predict reward structure in latent space. During deployment, the observation encoder and reward predictor are frozen while only the transition dynamics are updated via unsupervised rollouts, allowing the policy to adapt using imagined trajectories without new rewards.

By Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez