arXiv AI By Joery Ari\"en de Vries, Neil David Lawrence, Zhenwen Dai

Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT

Read the original on arXiv AI →

The paper introduces Follow the Winners (FTW), a critic‑free policy‑learning algorithm that adapts the cross‑entropy method for reinforcement fine‑tuning of agentic large language models. FTW replaces group rollouts with an ordinal filter on replay‑buffer samples, achieving polynomial concentration in the order statistic of returns and offering a bounded risk‑seeking offset that balances variance reduction. Experiments on Sokoban and Search‑R1 show that FTW matches the performance of GRPO and PPO while reducing reliance on value models or repeated rollouts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 20

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks.

arXiv Machine Learning
Sep 30

EasyPPO: Stabilizing the Critic Is Key

arXiv:2609.36802v1 Announce Type: new Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning...

By Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez