Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
arXiv:2608. 05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL).
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary prediction performed either in latent space (self-prediction) or observation space (observation prediction).
arXiv:2608. 05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL).
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
arXiv:2609.37250v1 Announce Type: cross Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.
Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.
The paper introduces Sampling-Guided Policy Search (SGPS), a method that combines sampling-based model‑predictive control with first‑order policy gradients to accelerate visual policy learning for locomotion and manipulation tasks. SGPS starts with behavior cloning from sampled actions and then alternates between sampling‑based refinement and short‑horizon policy updates under varied initial states and dynamics. The approach is demonstrated on simulated Unitree Go2 and G1 robots, learning tasks such as obstacle traversal and bimanual carrying, and the distilled policies transfer zero‑shot to a real Go2 robot using onboard depth perception.
The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
JEPA‑TTT is a method that continuously adapts the latent dynamics predictor of a pretrained Joint‑Embedding Predictive Architecture (JEPA) world model during test time. It performs self‑supervised updates across episodes while keeping the visual encoder and reward head fixed, using dense replay to sample prediction windows from a growing buffer. In experiments on eight dynamics shifts across four continuous‑control environments, JEPA‑TTT reduces latent prediction error by 83% and improves planning performance by 153% compared to the frozen model.
arXiv:2604. 16557v2 Announce Type: replace Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL).
arXiv:2607. 00796v1 Announce Type: new Abstract: Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks.
arXiv:2606. 12200v1 Announce Type: cross Abstract: We study policy representation learning from unlabeled multi-policy behavioral data.