Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
arXiv:2606. 05555v1 Announce Type: new Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge.
Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong performance, they rely on planning and complex training pipelines, making it unclear which components are essential for scalability.
arXiv:2606. 05555v1 Announce Type: new Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge.
arXiv:2602. 12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the representational strengths of model-based approaches, without incurring planning overhead.
arXiv:2605. 04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning.
arXiv:2602. 05031v2 Announce Type: replace Abstract: Planning with a learned model remains a key challenge in model-based reinforcement learning (RL).
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
arXiv:2602. 05999v3 Announce Type: replace Abstract: How does the amount of compute available to a reinforcement learning (RL) policy affect its learning?
arXiv:2608. 15509v1 Announce Type: cross Abstract: Task guided agents demonstrate strong performance in a wide range of complex tasks.
arXiv:2607. 20834v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets.
arXiv:2510. 14828v3 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully.
Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups.
DeepJEPA is a weight‑tied joint‑embedding predictive world model that treats transition depth as an inner test‑time scaling axis, learning when additional recurrent updates are worthwhile for each candidate and rollout step. Unlike traditional planners that uniformly deepen every transition, DeepJEPA concentrates extra computation on decision‑critical events such as contact onset and sustained object interaction, achieving comparable or better performance with only 1.00–1.26 updates per transition across five visual‑control settings. The approach demonstrates that improved planning does not require uniformly better object‑state decodability, but rather targeted internal computation where it can alter the planner’s elite set and action selection.
The Representation World Model (RWM) learns states, transitions, and executable plans directly within a representation space, bypassing traditional explicit dynamics models and action-space search. It uses inverse-dynamics supervision along latent paths to shape the representation geometry, enabling direct planning by constructing a latent path between current and goal states and recovering actions via inverse dynamics. Experiments on continuous-control benchmarks and robotic manipulation tasks demonstrate RWM’s effectiveness and potential for complex embodied control.