arXiv AI By Amir Esterhuysen, Anders Jonsson

The Terminal Representation in Reinforcement Learning

Read the original on arXiv AI →

arXiv:2605. 31289v2 Announce Type: replace-cross Abstract: Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Exploration and Online Transfer with Behavioral Foundation Models

arXiv:2606. 29980v1 Announce Type: new Abstract: Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories.

By Louis Bagot (SyCoSMA), Mathieu Lefort (LIRIS, SyCoSMA, IRISA, MALT, UR), La\"etitia Matignon (SyCoSMA)
arXiv Machine Learning
Sep 18

EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence

EmbodiedMind introduces a three-stage training paradigm for embodied foundation models that tackles inefficient sample use, task imbalance, and credit assignment in long-horizon planning. The stages—Rejection Sampling-based Fine‑Tuning, Iterative Rejection GRPO, and Trie‑GRPO—filter low‑informative data, balance task difficulty, and use action prefix trees for step‑level advantage estimation. This approach yields a state‑of‑the‑art average performance of 70.02% across 18 benchmarks, notably improving long‑horizon task planning accuracy.

By Feifan Wang, Zongbing Zhang, Yu Zhang, Lingfeng Wang, Yurui Zhu, Jin Deng, Mingliang Zhang, Zhengguang Gao, Yongcheng Wang, Jin Xu, Ri Yang
arXiv AI
Sep 25

PoEM: Predicting RL Outcomes from Existing Policies

PoEM predicts reinforcement learning outcomes for a new reward function using models already trained on other rewards. If the new reward is a linear combination of existing ones, the new policy’s log-space representation can be expressed as a linear combination of existing log-policies. Even when rewards are not linearly related, log-policies often span a low‑rank subspace, allowing the weighting coefficients to be estimated from reward or basis policy outputs, enabling policy approximation without additional RL training.

By Kimia Hamidieh, Giannis Daras, Antonio Torralba
arXiv Machine Learning
Aug 31

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

The paper introduces VICT, a method that leverages the internal structure of verifiable tasks to perform fine‑grained credit assignment for long‑horizon LLM agents. VICT exposes executable or evidence‑backed atoms from a task’s terminal verifier and traces them back to actions via dependency‑valid proof edges, redistributing advantage only along these edges. This approach improves performance on ALFWorld and WebShop compared to outcome‑only training and matches recent fine‑grained credit methods without requiring additional critics, labels, or inference‑time verifier access.

By Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma