arXiv AI By Jing Ma, Chenhao Dang, Mingjie Liao

AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining

Read the original on arXiv AI →

arXiv:2505. 23878v2 Announce Type: replace-cross Abstract: Optimizing pretraining data composition is pivotal for LLM generalization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?

The paper investigates how reusing past samples can improve the sample efficiency of Proximal Policy Optimization (PPO). Two variants, wPPO-U and wPPO-BH, are introduced within a multiple importance weighting framework, each reusing data from recent iterations while preserving core PPO mechanics. The authors derive theoretical policy improvement bounds for both variants and empirically evaluate their impact on continuous control tasks.

By Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli