arXiv AI By Xiaocan Li, Shiliang Wu, Zheng Shen

A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation

Read the original on arXiv AI →

arXiv:2512. 06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 30

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

arXiv:2606. 29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse.

By Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
arXiv AI
Sep 17

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

The paper identifies a failure mode in Proximal Policy Optimization (PPO) critics called Value Flattening, where state values vary sharply across intermediate states while critic predictions remain flat. The authors analyze this phenomenon theoretically and empirically, linking it to an implicit variance penalty and redundant updates from temporally correlated states. They propose Sparse Proximal Policy Optimization (SP³O), which applies the value loss to only a few well‑separated states, and demonstrate that this sparse supervision mitigates Value Flattening and improves policy performance on Qwen3‑Base across model sizes and evaluation suites.

By Yizhuo Li, Jianhao Yan, Yun Luo, Zhi Wang, Futing Wang, Rong-Xi Tan, Kanghui Tian, Ganqu Cui, Ning Ding, Peilin Zhao, Yafu Li, Yu Cheng
arXiv AI
2d ago

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo