arXiv AI By Bernd Frauenknecht, Emma Cramer, Artur Eisele, Paul Kruse, Lukas Kesper, Jonas Hertrampf, Ramil Sabirov, Jyotirmaya Patra, Johannes Berger, Paul Brunzema, Friedrich Solowjow, Sebastian Trimpe

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

Read the original on arXiv AI →

RLLBC-Lib is an educational code library designed to lower the entry barrier for students learning reinforcement learning (RL) in the context of learning-based control. It offers a comprehensive collection of tabular RL methods to reinforce theoretical foundations, followed by a deep RL library that mirrors the same design principles to highlight parallels between simple and state‑of‑the‑art approaches. The library also includes implementations that illustrate core RL principles, contrast RL with other learning‑based control methods, and serve as a foundation for programming assignments with automated grading.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 11

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

arXiv:2608. 07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly.

By Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, Clare Lyle
arXiv Machine Learning
1d ago

Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

The paper introduces Function-Structured Graph Reinforcement Learning (FSG‑RL), a framework that links subproblem graphs to Python code and uses multiple verifiers for feedback. It first trains a policy via supervised fine‑tuning to generate code from function graphs, then refines it with Group Relative Policy Optimization (GRPO) that employs answer‑gated rewards and span‑level credit assignment. On a benchmark combining GSM8K, MathQA, MATH, and Omni‑MATH, GRPO raises final‑answer accuracy from 43.25 % to 67.50 % and full‑solution success from 32.25 % to 52.25 %, with further improvements when teacher supervision is added.

By Zihan Liu, Xurong Xie