Performance-Driven Environment Abstraction with Multi-Timescale Learning
arXiv:2606. 17377v1 Announce Type: new Abstract: We study performance-driven environment abstraction for decision-making in large Markov decision processes.
arXiv:2606. 00427v1 Announce Type: new Abstract: State abstraction in reinforcement learning is usually formulated as a partition of states based on reward and transition similarity.
arXiv:2606. 17377v1 Announce Type: new Abstract: We study performance-driven environment abstraction for decision-making in large Markov decision processes.
The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.
The paper introduces topological necessities—mechanism‑invariant subgoals derived from the topology of successful trajectories—used to guide long‑horizon goal‑conditioned reinforcement learning. By computing homology in dimensions 0 and 1 over a transport‑weighted carrier, the authors obtain an enumerable gate set that forms a recursive topological gate hierarchy. These certified gates transfer across different embodiments (e.g., from PointMaze to Ant and Humanoid) without retraining, achieving state‑of‑the‑art performance on several benchmark tasks.
arXiv:2603. 08558v3 Announce Type: replace Abstract: Learning compact state representations in Markov Decision Processes (MDPs) has proven crucial for addressing the curse of dimensionality in large-scale reinforcement learning (RL) problems.
arXiv:2608. 06276v1 Announce Type: cross Abstract: Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure.
Spectral Prioritized Sweeping (SPS) extends traditional Prioritized Sweeping by incorporating graph topology through the resolvent and Laplacian diffusion, creating a smoother priority score that propagates reward changes more effectively in nonstationary reinforcement learning. The method, called Graph Topology Augmentation for Prioritized Sweeping (GTA-PS), uses a mixing of regularized Laplacian inverses and an adaptive scheduler based on the Second Largest Eigenvalue Modulus to adjust the influence of topology during replanning. Experiments on FourRooms and GARNET domains show that GTA-PS improves replanning efficiency compared to standard PS under both exact dynamic programming and Dyna-style planners.
arXiv:2607. 17038v1 Announce Type: new Abstract: This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sparse reward attribution, and dynamic environmental interaction, by designing and optimizing an intelligent agent workflow.
arXiv:2607. 19232v1 Announce Type: new Abstract: Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution.
arXiv:2609.14968v1 Announce Type: new Abstract: Online scheduling of dependency-aware tasks in heterogeneous cloud clusters is a fundamental yet challenging problem due to the complex interplay betwe...
The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.
arXiv:2304.10041v2 Announce Type: replace Abstract: This work investigates formal policy synthesis for continuous-state stochastic dynamic systems subject to high-level specifications expressed in li...
The paper introduces Partial GFlowNet, a method that partitions a large state space into overlapping partial state spaces to accelerate convergence of Generative Flow Networks. By restricting the actor’s exploration to these smaller regions and using a heuristic to switch between them, the approach enables efficient identification of high‑reward subregions. Experiments on popular datasets show that Partial GFlowNet converges faster, produces higher‑reward candidates, and improves diversity compared to existing methods.