Task-Relevant Representation Decoupling for Visual Reinforcement Learning Generalization
arXiv:2607. 00796v1 Announce Type: new Abstract: Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks.
arXiv:2606. 11860v1 Announce Type: new Abstract: In this paper, we introduce Representation Prediction via Autoencoding using Iterative Refinement (RePAIR) - a novel self-supervised representation learning architecture that synthesizes Masked Autoencoders (MAE), Joint Embedding Predictive Architectures (JEPA), and Bidirectional Encoder Representations from Transformers (BERT).
arXiv:2607. 00796v1 Announce Type: new Abstract: Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks.
arXiv:2501. 14622v5 Announce Type: replace Abstract: Learning efficient representations for decision-making policies is a challenge in imitation learning (IL).
arXiv:2607. 00190v1 Announce Type: cross Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games.
arXiv:2607. 04153v1 Announce Type: cross Abstract: Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information.
arXiv:2511. 05963v4 Announce Type: replace Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc lookups over past tokens.
The paper introduces Skill Abstraction with Interpretable Latents (SAIL), a method that models human skill as a persistent, multi‑dimensional construct inferred from naturalistic behavior over time. SAIL produces a robust skill embedding that blends expert and novice bases, learns transferable subskills through counterfactual subskill swaps, and supports skill‑informed behavior prediction across various in‑domain contexts. Experiments on racing and baseball demonstrate that SAIL achieves strong predictive performance, improves behaviorally grounded disentanglement compared to baselines, and enhances downstream AI coaching outcomes.
arXiv:2606. 12200v1 Announce Type: cross Abstract: We study policy representation learning from unlabeled multi-policy behavioral data.
arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
arXiv:2609.23881v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction....
arXiv:2606. 09327v1 Announce Type: cross Abstract: Football event data constitute a rich spatiotemporal source for quantitative analysis of player actions in team sports.
arXiv:2606. 14765v1 Announce Type: cross Abstract: Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning.
arXiv:2603. 12231v2 Announce Type: replace Abstract: Learning good representations is essential for latent planning with world models.