Gradient-based Planning for World Models at Longer Horizons
grasp-results-table table { font-size: 0. 875rem; line-height: 1.
In this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks.
grasp-results-table table { font-size: 0. 875rem; line-height: 1.
arXiv:2606. 29980v1 Announce Type: new Abstract: Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories.
arXiv:2602. 00781v2 Announce Type: replace Abstract: Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time.
What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task.
arXiv:2609.36393v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit...
PoEM predicts reinforcement learning outcomes for a new reward function using models already trained on other rewards. If the new reward is a linear combination of existing ones, the new policy’s log-space representation can be expressed as a linear combination of existing log-policies. Even when rewards are not linearly related, log-policies often span a low‑rank subspace, allowing the weighting coefficients to be estimated from reward or basis policy outputs, enabling policy approximation without additional RL training.
Posted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-following framework. This paradigm is exemplified by recent multi-task LLMs, such as T0 , FLAN , and OPT-IML .
arXiv:2606. 18812v1 Announce Type: cross Abstract: Foundation models for language and vision are powered by internet-scale data, while structured domains (tabular prediction, time-series forecasting, graph learning, reinforcement learning) are not.
arXiv:2609.40149v1 Announce Type: new Abstract: Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can dif...
apr-fig { text-align: center; margin: 1. 35em 0; line-height: 1.
arXiv:2610.00592v1 Announce Type: cross Abstract: In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next...
Foundation models for language and vision are powered by internet-scale data, while structured domains (tabular prediction, time-series forecasting, graph learning, reinforcement learning) are not. The substitute is synthetic data, which shifts the burden from collection to prior design.