arXiv Machine Learning By Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2601. 23075v2 Announce Type: replace Abstract: On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy, and policy updates must be conservative.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 28

Simple Actors and Deep Critics for Scalable Reinforcement Learning

The paper introduces LAC (Light Actor, deep Critic), an offline reinforcement learning approach that allocates model capacity to a deep critic rather than a complex actor to improve inference efficiency. It addresses three failure modes—optimization, bootstrap-noise amplification, and value-range drift—using a residual MLP backbone, n‑step bootstrap targets, and a categorical cross‑entropy loss. Experiments on OGBench show LAC matches state‑of‑the‑art diffusion and flow‑matching baselines while reducing inference latency by up to four times.

By Guhyeon Kang, Jaehwi Lee, Minhae Kwon
arXiv Machine Learning
Jul 21

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

arXiv:2509. 02522v3 Announce Type: replace-cross Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have empowered large language models (LLMs) to tackle challenging reasoning tasks such as mathematics and programming, however existing RLVR methods often suffer from sparse reward signals and unstable policy gradient updates inherent to RL-based approaches.

By Jiaming Li, Longze Chen, Ze Gong, Yukun Chen, Lu Wang, Wanwei He, Run Luo, Min Yang