grasp-results-table table { font-size: 0. 875rem; line-height: 1.
arXiv:2606. 29980v1 Announce Type: new Abstract: Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories.
By Louis Bagot (SyCoSMA), Mathieu Lefort (LIRIS, SyCoSMA, IRISA, MALT, UR), La\"etitia Matignon (SyCoSMA)
arXiv:2602. 00781v2 Announce Type: replace Abstract: Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time.
By Jiamin Xu, Kyra Gan
What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task.
arXiv:2609.36393v1 Announce Type: cross
Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit...
By Muhang Tian, Sherry Yang
PoEM predicts reinforcement learning outcomes for a new reward function using models already trained on other rewards. If the new reward is a linear combination of existing ones, the new policy’s log-space representation can be expressed as a linear combination of existing log-policies. Even when rewards are not linearly related, log-policies often span a low‑rank subspace, allowing the weighting coefficients to be estimated from reward or basis policy outputs, enabling policy approximation without additional RL training.
By Kimia Hamidieh, Giannis Daras, Antonio Torralba