Hugging Face Trending Papers

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

arXiv AI
Sep 17

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

RLLBC-Lib is an educational code library designed to lower the entry barrier for students learning reinforcement learning (RL) in the context of learning-based control. It offers a comprehensive collection of tabular RL methods to reinforce theoretical foundations, followed by a deep RL library that mirrors the same design principles to highlight parallels between simple and state‑of‑the‑art approaches. The library also includes implementations that illustrate core RL principles, contrast RL with other learning‑based control methods, and serve as a foundation for programming assignments with automated grading.

By Bernd Frauenknecht, Emma Cramer, Artur Eisele, Paul Kruse, Lukas Kesper, Jonas Hertrampf, Ramil Sabirov, Jyotirmaya Patra, Johannes Berger, Paul Brunzema, Friedrich Solowjow, Sebastian Trimpe
arXiv Machine Learning
Aug 11

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

arXiv:2608. 07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly.

By Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, Clare Lyle
arXiv AI
Aug 19

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

EXPO-FT is a system that enables stable, sample‑efficient reinforcement learning fine‑tuning of pretrained Vision‑Language‑Action (VLA) policies. It achieves perfect success on a range of manipulation tasks—such as routing string lights, striking a pool ball, and inserting a flower into a wine bottle—using only about 19.1 minutes of online robot data. The approach outperforms both RL-from-scratch and existing VLA fine‑tuning methods, and the authors provide an open‑source codebase to support wider adoption.

By Perry Dong, Kuo-Han Hung, Tian Gao, Dorsa Sadigh, Chelsea Finn
arXiv Machine Learning
1d ago

Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

The paper introduces Function-Structured Graph Reinforcement Learning (FSG‑RL), a framework that links subproblem graphs to Python code and uses multiple verifiers for feedback. It first trains a policy via supervised fine‑tuning to generate code from function graphs, then refines it with Group Relative Policy Optimization (GRPO) that employs answer‑gated rewards and span‑level credit assignment. On a benchmark combining GSM8K, MathQA, MATH, and Omni‑MATH, GRPO raises final‑answer accuracy from 43.25 % to 67.50 % and full‑solution success from 32.25 % to 52.25 %, with further improvements when teacher supervision is added.

By Zihan Liu, Xurong Xie