arXiv AI

Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT

The paper introduces Follow the Winners (FTW), a critic‑free policy‑learning algorithm that adapts the cross‑entropy method for reinforcement fine‑tuning of agentic large language models. FTW replaces group rollouts with an ordinal filter on replay‑buffer samples, achieving polynomial concentration in the order statistic of returns and offering a bounded risk‑seeking offset that balances variance reduction. Experiments on Sokoban and Search‑R1 show that FTW matches the performance of GRPO and PPO while reducing reliance on value models or repeated rollouts.

Hugging Face Trending Papers
Aug 20

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks.

arXiv Machine Learning
Sep 30

EasyPPO: Stabilizing the Critic Is Key

arXiv:2609.36802v1 Announce Type: new Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning...

By Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez
arXiv AI
Oct 2

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo
arXiv AI
22h ago

Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning

The paper introduces Entropy‑Normalized Trust Region (ENTR), a method for asynchronous reinforcement learning that adjusts off‑policy ratio thresholds based on token entropy rather than a single magnitude cut‑off. By recognizing that the natural scale of the ratio is set by entropy, ENTR preserves genuine exploration while filtering out noise from low‑entropy, stale data. Experiments on long‑horizon agentic tasks and mathematical reasoning benchmarks show ENTR outperforms existing asynchronous methods, improving BrowseComp‑Plus performance by 6.9 % and enabling stable training up to 30 policy versions of staleness while matching synchronous GRPO at a 2.6× speedup.

By Guanqun Zhao, Zijun Xie, Binbin Zheng, Yehan Yang, Jiafeng Lu, Aoqi Hu, Enlei Gong, Zeyu Chen