arXiv Machine Learning

Simple Actors and Deep Critics for Scalable Reinforcement Learning

The paper introduces LAC (Light Actor, deep Critic), an offline reinforcement learning approach that allocates model capacity to a deep critic rather than a complex actor to improve inference efficiency. It addresses three failure modes—optimization, bootstrap-noise amplification, and value-range drift—using a residual MLP backbone, n‑step bootstrap targets, and a categorical cross‑entropy loss. Experiments on OGBench show LAC matches state‑of‑the‑art diffusion and flow‑matching baselines while reducing inference latency by up to four times.

arXiv Machine Learning
Jun 25

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

arXiv:2601. 23075v2 Announce Type: replace Abstract: On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy, and policy updates must be conservative.

By Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi
arXiv AI
Jun 16

QPILOTS: Efficient Test-Time Q-Steering for Flow Policies

arXiv:2606. 14801v1 Announce Type: cross Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult.

By Yifan Ruan, Chenyang Cao, Andreas Burger, Ali Pesaranghader, Kaveh Kamali, Jaehong Kim, Nandita Vijaykumar, Alan Aspuru-Guzik, Igor Gilitschenski, Nicholas Rhinehart
arXiv AI
Jul 21

MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models

arXiv:2607. 18006v1 Announce Type: cross Abstract: Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($\leq 4 \, \mathrm{B}$ parameters) trained under limited budgets.

By Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Zifeng Ding, Volker Tresp, Yunpu Ma
arXiv AI
3d ago

Trust the Critic More

The paper introduces Actor‑Critic with Action Chunking (AC2), a method that assigns credit to short action chunks instead of entire trajectories, enabling policy updates without waiting for terminal rewards. AC2 employs local readiness, reference solutions, and 10k‑token chunks to make critic‑based credit assignment reliable. Experiments on Qwen3‑4B with FineProofs‑RL show AC2 surpasses GRPO’s peak validation score while using 2.5× fewer decoding FLOPs and fewer training steps.

By Kaiyue Wen, Luke Bailey, Arvind Mahankali, Tengyu Ma