Hugging Face Trending Papers

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

Read the original on Hugging Face Trending Papers →

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.

OpenAI Blog
Aug 3, 2017

Gathering human feedback

RL-Teacher is an open-source implementation of our interface to train AIs via occasional human feedback rather than hand-crafted reward functions. The underlying technique was developed as a step towards safe AI systems, but also applies to reinforcement learning problems with rewards that are hard to specify.