arXiv AI By Guanqun Zhao, Zijun Xie, Binbin Zheng, Yehan Yang, Jiafeng Lu, Aoqi Hu, Enlei Gong, Zeyu Chen

Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning

Read the original on arXiv AI →

The paper introduces Entropy‑Normalized Trust Region (ENTR), a method for asynchronous reinforcement learning that adjusts off‑policy ratio thresholds based on token entropy rather than a single magnitude cut‑off. By recognizing that the natural scale of the ratio is set by entropy, ENTR preserves genuine exploration while filtering out noise from low‑entropy, stale data. Experiments on long‑horizon agentic tasks and mathematical reasoning benchmarks show ENTR outperforms existing asynchronous methods, improving BrowseComp‑Plus performance by 6.9 % and enabling stable training up to 30 policy versions of staleness while matching synchronous GRPO at a 2.6× speedup.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 3

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

arXiv:2607. 22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse.

By Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
arXiv Machine Learning
Jun 5

Extreme Region Policy Distillation

arXiv:2605. 25582v2 Announce Type: replace Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces distribution mismatch that existing trust-region techniques mitigate primarily by enforcing conservative optimization, often leaving rich training signals underutilized.

By Changyu Chen, Xiting Wang, Rui Yan