arXiv AI By Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Read the original on arXiv AI →

The paper introduces the Posterior Concentration Phenomenon (PCP), a length‑dependent failure mode where probability‑based rewards collapse to a narrow interval for long reasoning traces, destabilizing verifier‑free reinforcement learning. To address this, the authors propose RLCPR, a framework that uses uncertainty‑aware data sampling and concentration‑aware regularization to mitigate PCP, improving token efficiency and outperforming state‑of‑the‑art baselines on multiple reasoning benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo