arXiv:2606. 17377v1 Announce Type: new Abstract: We study performance-driven environment abstraction for decision-making in large Markov decision processes.
By Yue Guan, Dipankar Maity, Panagiotis Tsiotras
The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.
By Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup
The paper introduces topological necessities—mechanism‑invariant subgoals derived from the topology of successful trajectories—used to guide long‑horizon goal‑conditioned reinforcement learning. By computing homology in dimensions 0 and 1 over a transport‑weighted carrier, the authors obtain an enumerable gate set that forms a recursive topological gate hierarchy. These certified gates transfer across different embodiments (e.g., from PointMaze to Ant and Humanoid) without retraining, achieving state‑of‑the‑art performance on several benchmark tasks.
By Hao Shi, Xi Li
arXiv:2603. 08558v3 Announce Type: replace Abstract: Learning compact state representations in Markov Decision Processes (MDPs) has proven crucial for addressing the curse of dimensionality in large-scale reinforcement learning (RL) problems.
By Tommaso Giorgi, Pierriccardo Olivieri, Keyue Jiang, Laura Toni, Matteo Papini
arXiv:2608. 06276v1 Announce Type: cross Abstract: Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure.
By Farzana Nasrin
Spectral Prioritized Sweeping (SPS) extends traditional Prioritized Sweeping by incorporating graph topology through the resolvent and Laplacian diffusion, creating a smoother priority score that propagates reward changes more effectively in nonstationary reinforcement learning. The method, called Graph Topology Augmentation for Prioritized Sweeping (GTA-PS), uses a mixing of regularized Laplacian inverses and an adaptive scheduler based on the Second Largest Eigenvalue Modulus to adjust the influence of topology during replanning. Experiments on FourRooms and GARNET domains show that GTA-PS improves replanning efficiency compared to standard PS under both exact dynamic programming and Dyna-style planners.
By Hung Pham, Tuan Dam