arXiv Machine Learning By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

Topology-Aware State Abstraction with Tangle Cores for Markov Decision Processes

Read the original on arXiv Machine Learning →

arXiv:2606. 00427v1 Announce Type: new Abstract: State abstraction in reinforcement learning is usually formulated as a partition of states based on reward and transition similarity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.

By Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup
arXiv Machine Learning
Sep 11

Topological Necessities: Mechanism-Invariant Strategic Subgoals for Cross-Embodiment Goal-Conditioned Control

The paper introduces topological necessities—mechanism‑invariant subgoals derived from the topology of successful trajectories—used to guide long‑horizon goal‑conditioned reinforcement learning. By computing homology in dimensions 0 and 1 over a transport‑weighted carrier, the authors obtain an enumerable gate set that forms a recursive topological gate hierarchy. These certified gates transfer across different embodiments (e.g., from PointMaze to Ant and Humanoid) without retraining, achieving state‑of‑the‑art performance on several benchmark tasks.

By Hao Shi, Xi Li
arXiv Machine Learning
Sep 10

Spectral Prioritized Sweeping in Nonstationary Reinforcement Learning

Spectral Prioritized Sweeping (SPS) extends traditional Prioritized Sweeping by incorporating graph topology through the resolvent and Laplacian diffusion, creating a smoother priority score that propagates reward changes more effectively in nonstationary reinforcement learning. The method, called Graph Topology Augmentation for Prioritized Sweeping (GTA-PS), uses a mixing of regularized Laplacian inverses and an adaptive scheduler based on the Second Largest Eigenvalue Modulus to adjust the influence of topology during replanning. Experiments on FourRooms and GARNET domains show that GTA-PS improves replanning efficiency compared to standard PS under both exact dynamic programming and Dyna-style planners.

By Hung Pham, Tuan Dam