← Back to all news
OpenAI Blog April 10, 2018

Gotta Learn Fast: A new benchmark for generalization in RL

Read the original on OpenAI Blog →

The Flow has not summarised this story yet — read it at OpenAI Blog.

  • reinforcement-learning
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 2

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

arXiv:2510. 10541v2 Announce Type: replace-cross Abstract: Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).

By Zihan Chen, Yiming Zhang, Hengguang Zhou, Zenghui Ding, Yining Sun, Cho-Jui Hsieh
llmsreinforcement-learningbenchmarks
More like this →
OpenAI Blog
Nov 9, 2016

RL²: Fast reinforcement learning via slow reinforcement learning

reinforcement-learning
More like this →
arXiv AI
Jun 2

Certificate-Guided Evaluation of Reinforcement Learning Generalization

arXiv:2606. 00840v1 Announce Type: new Abstract: This work presents a logic-driven framework to evaluate the performance of reinforcement learning (RL) algorithms in their ability to generalize to unseen tasks.

By Vignesh Subramanian, {\DJ}or{\dj}e \v{Z}ikeli\'c, Suguman Bansal
reinforcement-learningbenchmarks
More like this →
arXiv AI
Jul 1

Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR

arXiv:2606. 31813v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tuning (SFT) paradigm.

By Ruijia Zhang, Jiacheng Zhu, Hanqing Zhu, Laixi Shi
llmsreinforcement-learningfine-tuningbenchmarks
More like this →
arXiv AI
Sep 15

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

arXiv:2609.13443v1 Announce Type: cross Abstract: We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvements on easy problems that an...

By Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville
llmsreinforcement-learningbenchmarks
More like this →
Hugging Face Blog
Oct 24, 2023

Exploring simple optimizations for SDXL

More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea