OpenAI Baselines: ACKTR & A2C
We’re releasing two new OpenAI Baselines implementations: ACKTR and A2C. A2C is a synchronous, deterministic variant of Asynchronous Advantage Actor Critic (A3C) which we’ve found gives equal performance.
We’re releasing two new OpenAI Baselines implementations: ACKTR and A2C. A2C is a synchronous, deterministic variant of Asynchronous Advantage Actor Critic (A3C) which we’ve found gives equal performance.
arXiv:2609.36058v1 Announce Type: new Abstract: Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient method...
arXiv:2607. 23605v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns.
arXiv:2602. 17632v3 Announce Type: replace-cross Abstract: Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with value-based RL algorithms typically causes immediate drops in performance.
arXiv:2512. 05291v3 Announce Type: replace Abstract: Actor-critic (AC) methods are a cornerstone of reinforcement learning (RL) but offer limited interpretability.
arXiv:2509. 26000v3 Announce Type: replace Abstract: Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability.
arXiv:1908.08773v3 Announce Type: replace Abstract: In certain reinforcement learning (RL) scenarios there are adversaries trying to interfere with the underlying reward process for their own benefit...
arXiv:2609.26355v1 Announce Type: new Abstract: Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted ma...
We’re launching a transfer learning contest that measures a reinforcement learning algorithm’s ability to generalize from previous experience.
DAMPER is a new technique for actor‑critic methods that reduces action oscillation in continuous control tasks. It combines the native actor gradient with a temporal‑consistency gradient using conflict‑conditioned projection and adaptive magnitude control, ensuring the auxiliary component aligns positively with the actor gradient. Experiments on TD3 and SAC across six tasks show that DAMPER consistently lowers oscillation compared to baseline agents and outperforms other methods in most task‑backbone pairs.
The paper introduces Best‑Practice Critic Optimization (BPCO), a stable and efficient recipe for training a critic in reinforcement learning for large language models. BPCO combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized policy advantages, and length‑adaptive advantage estimation, allowing the critic to be conditioned on hidden reward information. Experiments on mathematical reasoning tasks with models from 1.5B to 30B parameters show that BPCO consistently outperforms a strong critic‑based baseline and matches or exceeds group‑based methods while sampling only one response per prompt.