Berkeley AI Research

RL without TD learning

In this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks.

arXiv AI
Jun 30

Exploration and Online Transfer with Behavioral Foundation Models

arXiv:2606. 29980v1 Announce Type: new Abstract: Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories.

By Louis Bagot (SyCoSMA), Mathieu Lefort (LIRIS, SyCoSMA, IRISA, MALT, UR), La\"etitia Matignon (SyCoSMA)
arXiv AI
Sep 25

PoEM: Predicting RL Outcomes from Existing Policies

PoEM predicts reinforcement learning outcomes for a new reward function using models already trained on other rewards. If the new reward is a linear combination of existing ones, the new policy’s log-space representation can be expressed as a linear combination of existing log-policies. Even when rewards are not linearly related, log-policies often span a low‑rank subspace, allowing the weighting coefficients to be estimated from reward or basis policy outputs, enabling policy approximation without additional RL training.

By Kimia Hamidieh, Giannis Daras, Antonio Torralba
Google AI Blog
Mar 14, 2024

Cappy: Outperforming and boosting large multi-task language models with a small scorer

Posted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-following framework. This paradigm is exemplified by recent multi-task LLMs, such as T0 , FLAN , and OPT-IML .

By Google AI