Policy Gradient with PyTorch
Related stories
Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning
arXiv:2609.06882v1 Announce Type: cross Abstract: Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remain...
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
arXiv:2605. 18591v2 Announce Type: replace Abstract: Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the Fisher matrix.
Equivalence between policy gradients and soft Q-learning
Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
arXiv:2605. 14982v2 Announce Type: replace-cross Abstract: We address the discounted reward setting in reinforcement learning (RL).
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
arXiv:2605. 04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning.
Proximal Policy Optimization (PPO)
Proximal Policy Optimization for Amortized Discrete Sampling
arXiv:2606. 15793v1 Announce Type: cross Abstract: This paper explores policy gradient algorithms for training stochastic policies to sample from structured discrete probability distributions under the Generative Flow Network (GFlowNet) framework.
Lipschitz-Regularized Critics Lead to Policy Robustness Against Transition Dynamics Uncertainty
arXiv:2404. 13879v5 Announce Type: replace Abstract: Uncertainties in transition dynamics pose a critical challenge in reinforcement learning (RL), often resulting in performance degradation of trained policies when deployed on hardware.
Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization
The paper introduces a geometric framework for reinforcement learning that treats policies as mappings into the Wasserstein space of action probabilities. It establishes a Riemannian structure induced by stationary distributions, defines the tangent space of policies, and characterizes geodesics while addressing measurability concerns. The authors formulate a general RL optimization problem, construct a gradient flow via Otto's calculus, compute the gradient and Hessian of the energy, and demonstrate the approach with numerical examples for low‑dimensional problems and neural‑network‑parameterized policies for high‑dimensional settings.
Evolved Policy Gradients
We’re releasing an experimental metalearning approach called Evolved Policy Gradients, a method that evolves the loss function of learning agents, which can enable fast training on novel tasks. Agents trained with EPG can succeed at basic tasks at test time that were outside their training regime, like learning to navigate to an object on a different side of the room from where it was placed during training.