arXiv AI

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

RLLBC-Lib is an educational code library designed to lower the entry barrier for students learning reinforcement learning (RL) in the context of learning-based control. It offers a comprehensive collection of tabular RL methods to reinforce theoretical foundations, followed by a deep RL library that mirrors the same design principles to highlight parallels between simple and state‑of‑the‑art approaches. The library also includes implementations that illustrate core RL principles, contrast RL with other learning‑based control methods, and serve as a foundation for programming assignments with automated grading.

arXiv Machine Learning
Aug 11

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

arXiv:2608. 07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly.

By Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, Clare Lyle
arXiv Machine Learning
1d ago

Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

The paper introduces Function-Structured Graph Reinforcement Learning (FSG‑RL), a framework that links subproblem graphs to Python code and uses multiple verifiers for feedback. It first trains a policy via supervised fine‑tuning to generate code from function graphs, then refines it with Group Relative Policy Optimization (GRPO) that employs answer‑gated rewards and span‑level credit assignment. On a benchmark combining GSM8K, MathQA, MATH, and Omni‑MATH, GRPO raises final‑answer accuracy from 43.25 % to 67.50 % and full‑solution success from 32.25 % to 52.25 %, with further improvements when teacher supervision is added.

By Zihan Liu, Xurong Xie
arXiv Machine Learning
Jul 21

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

arXiv:2509. 02522v3 Announce Type: replace-cross Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have empowered large language models (LLMs) to tackle challenging reasoning tasks such as mathematics and programming, however existing RLVR methods often suffer from sparse reward signals and unstable policy gradient updates inherent to RL-based approaches.

By Jiaming Li, Longze Chen, Ze Gong, Yukun Chen, Lu Wang, Wanwei He, Run Luo, Min Yang