Hugging Face Blog

Advantage Actor Critic (A2C)

arXiv AI
3d ago

DAMPER: Return-Prioritized Gradient Control for Smooth Policies

DAMPER is a new technique for actor‑critic methods that reduces action oscillation in continuous control tasks. It combines the native actor gradient with a temporal‑consistency gradient using conflict‑conditioned projection and adaptive magnitude control, ensuring the auxiliary component aligns positively with the actor gradient. Experiments on TD3 and SAC across six tasks show that DAMPER consistently lowers oscillation compared to baseline agents and outperforms other methods in most task‑backbone pairs.

By Seokmin Ko, Taewon Goo, Kihyuk Hong
arXiv AI
Aug 25

How to Train a Critic Stably and Efficiently

The paper introduces Best‑Practice Critic Optimization (BPCO), a stable and efficient recipe for training a critic in reinforcement learning for large language models. BPCO combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized policy advantages, and length‑adaptive advantage estimation, allowing the critic to be conditioned on hidden reward information. Experiments on mathematical reasoning tasks with models from 1.5B to 30B parameters show that BPCO consistently outperforms a strong critic‑based baseline and matches or exceeds group‑based methods while sampling only one response per prompt.

By Penghui Qi, Xiangxin Zhou, Wee Sun Lee