arXiv AI By Ting Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen, Daoyuan Chen

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 26917v1 Announce Type: cross Abstract: Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 30

EasyPPO: Stabilizing the Critic Is Key

arXiv:2609.36802v1 Announce Type: new Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning...

By Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez
arXiv Machine Learning
Jun 5

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter

arXiv:2606. 05800v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals.

By Powei Chang, Jinpeng Zhang, Chaoqun Sun, MiniWell Tsao, Lianrui Li, Jianxiang Xiang, Chenyu Wang, Yukang Gao, Dongying Kong