Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
arXiv:2605. 14982v2 Announce Type: replace-cross Abstract: We address the discounted reward setting in reinforcement learning (RL).
arXiv:2512. 05291v3 Announce Type: replace Abstract: Actor-critic (AC) methods are a cornerstone of reinforcement learning (RL) but offer limited interpretability.
arXiv:2605. 14982v2 Announce Type: replace-cross Abstract: We address the discounted reward setting in reinforcement learning (RL).
arXiv:2609.36058v1 Announce Type: new Abstract: Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient method...
arXiv:2608.20909v1 Announce Type: new Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled l...
Limiting‑Kernel Q(λ) (LKQL) is an off‑policy value estimator that blends n‑step truncation with a long‑horizon approximation based on the limiting kernel. It maintains the computational efficiency of n‑step methods while improving policy evaluation accuracy, especially for long‑horizon tasks. The authors prove faster convergence of LKQL’s operator under aperiodicity and near‑on‑policy conditions, and demonstrate empirical gains on MuJoCo continuous‑control benchmarks.
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations rel...
arXiv:2606. 20008v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become a central tool for improving the reasoning ability of large language models, but current methods face a trade-off between simplicity and credit assignment.
arXiv:2607. 01880v1 Announce Type: new Abstract: Value functions are an essential component in actor-critic based deep reinforcement learning (RL).
arXiv:2606. 25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints.
arXiv:2510. 19244v3 Announce Type: replace Abstract: Deep reinforcement learning (RL) achieves remarkable performance but lacks interpretability, limiting trust in policy behavior.
arXiv:2507. 04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learning, and Actor-Critic methods.
The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.
arXiv:2506. 13862v2 Announce Type: replace-cross Abstract: In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive policies has emerged as a popular tool both in theory and practice.