arXiv AI By Anton Bolychev, Georgiy Malaniya, Sinan Ibrahim, Pavel Osinenko

An Agency-Transferring Model-Free Policy Enhancement Technique

Read the original on arXiv AI →

arXiv:2606. 09825v1 Announce Type: cross Abstract: Training reinforcement learning (RL) policies from scratch is costly: it requires careful reward and environment design, extensive tuning, and substantial computation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

MInTRL: Off-policy Intervention can boost On-policy RL

MInTRL (Minimal Intervention Reinforcement Learning) expands exploration in on-policy reinforcement learning by inserting sparse, local corrections into rollouts via a judge-intervention policy. These interventions replace erroneous suffixes and immediately return control to the main policy, allowing the agent to explore beyond its natural trajectory while maintaining on-policy data. The method uses a sequence-level advantage-regression objective, avoiding importance sampling, and demonstrates superior performance on math and code benchmarks compared to standard on-policy and off-policy baselines.

By Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong
arXiv Machine Learning
23h ago

Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?

The paper investigates how reusing past samples can improve the sample efficiency of Proximal Policy Optimization (PPO). Two variants, wPPO-U and wPPO-BH, are introduced within a multiple importance weighting framework, each reusing data from recent iterations while preserving core PPO mechanics. The authors derive theoretical policy improvement bounds for both variants and empirically evaluate their impact on continuous control tasks.

By Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli