← Back to all news
arXiv Machine Learning September 15, 2026 By Qiming Dai, Yin Liu, Junyu Zhang, Zaiwen Wen

Non-Asymptotic Global Convergence of PPO-Clip

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 9

Rethinking the Divergence Regularization in LLM RL

arXiv:2606. 09821v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a key component of post-training large language models (LLMs).

By Jiarui Yao, Xiangxin Zhou, Penghui Qi, Wee Sun Lee, Liefeng Bo, Tianyu Pang
llmsreinforcement-learning
More like this →
arXiv AI
Jun 15

Rethinking the Trust Region in LLM Reinforcement Learning

arXiv:2602. 04879v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorithm.

By Penghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang, Chao Du, Min Lin, Wee Sun Lee
llmsreinforcement-learningfine-tuningefficiency
More like this →
arXiv Machine Learning
Jul 9

Lipschitz-Regularized Critics Lead to Policy Robustness Against Transition Dynamics Uncertainty

arXiv:2404. 13879v5 Announce Type: replace Abstract: Uncertainties in transition dynamics pose a critical challenge in reinforcement learning (RL), often resulting in performance degradation of trained policies when deployed on hardware.

By Xulin Chen, Ruipeng Liu, Zhenyu Gan, Garrett E. Katz
reinforcement-learningroboticssafety
More like this →
arXiv Machine Learning
Jun 19

VIMPO: Value-Implicit Policy Optimization for LLMs

arXiv:2606. 20008v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become a central tool for improving the reasoning ability of large language models, but current methods face a trade-off between simplicity and credit assignment.

By Zhewei Kang, Aosong Feng, Sergey Levine, Dawn Song, Xuandong Zhao
llmsreinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jun 2

Stabilizing Policy Optimization via Logits Convexity

arXiv:2603. 00963v2 Announce Type: replace Abstract: While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to supervised fine-tuning (SFT).

By Hongzhan Chen, Tao Yang, Yuhua Zhu, Shiping Gao, Xiaojun Quan, Ting Yao
llmsreinforcement-learningfine-tuningbenchmarks
More like this →
arXiv AI
Jul 21

Symmetric Behavior Regularized Policy Optimization

arXiv:2508. 04225v4 Announce Type: replace-cross Abstract: Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning.

By Lingwei Zhu, Haseeb Shah, Zheng Chen, Martha White
reinforcement-learningbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea