EasyPPO: Stabilizing the Critic Is Key
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 02181v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets.
arXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
The paper identifies a failure mode in Proximal Policy Optimization (PPO) critics called Value Flattening, where state values vary sharply across intermediate states while critic predictions remain flat. The authors analyze this phenomenon theoretically and empirically, linking it to an implicit variance penalty and redundant updates from temporally correlated states. They propose Sparse Proximal Policy Optimization (SP³O), which applies the value loss to only a few well‑separated states, and demonstrate that this sparse supervision mitigates Value Flattening and improves policy performance on Qwen3‑Base across model sizes and evaluation suites.
arXiv:2609.26355v1 Announce Type: new Abstract: Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted ma...
arXiv:2607. 25091v1 Announce Type: new Abstract: The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated.
arXiv:2608. 19842v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models.