Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries
Related stories
The Open Source Community is backing OpenEnv for Agentic RL
Introducing gpt-oss
We’re releasing gpt-oss-120b and gpt-oss-20b—two state-of-the-art open-weight language models that deliver strong real-world performance at low cost. Available under the flexible Apache 2.
OckBench: Measuring the Efficiency of LLM Reasoning
arXiv:2511. 05722v3 Announce Type: replace-cross Abstract: Large language models (LLMs) such as GPT-5 and Gemini 3 have pushed the frontier of automated reasoning and code generation.
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
Reinforcement learning (RL) is an exciting concept as well as a remarkable success story worth sharing. However, RL builds on rather complex interactions between different objects that play out over s...
LLM4RTL: Tool-Assisted LLM for RTL Generation
arXiv:2606. 15500v1 Announce Type: cross Abstract: Large language models (LLMs) have facilitated impressive progress in software engineering, code generation, tooling, and systems.
Introducing GPT-5.4
Introducing GPT-5. 4, OpenAI’s most most capable and efficient frontier model for professional work, with state-of-the-art coding, computer use, tool search, and 1M-token context.
RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
RLLBC-Lib is an educational code library designed to lower the entry barrier for students learning reinforcement learning (RL) in the context of learning-based control. It offers a comprehensive collection of tabular RL methods to reinforce theoretical foundations, followed by a deep RL library that mirrors the same design principles to highlight parallels between simple and state‑of‑the‑art approaches. The library also includes implementations that illustrate core RL principles, contrast RL with other learning‑based control methods, and serve as a foundation for programming assignments with automated grading.
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
arXiv:2511. 07317v2 Announce Type: replace-cross Abstract: We introduce Reinforcement Learning (RL) with Adaptive Verifiable Environments (RLVE), an approach using verifiable environments that procedurally generate problems and provide algorithmically verifiable rewards, to scale up RL for language models (LMs).
TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models
arXiv:2607. 01235v1 Announce Type: cross Abstract: Understanding how Large Language Models (LLMs) make token-level decisions during code generation remains a major challenge for both researchers and practitioners.
Introducing GPT-5.2-Codex
GPT-5. 2-Codex is OpenAI’s most advanced coding model, offering long-horizon reasoning, large-scale code transformations, and enhanced cybersecurity capabilities.
XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs
arXiv:2601. 04426v4 Announce Type: replace Abstract: Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols.