arXiv Machine Learning By Aleksandar Vujinovic, Aleksandar Kovacevic

ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning

Read the original on arXiv Machine Learning →

arXiv:2501. 14622v5 Announce Type: replace Abstract: Learning efficient representations for decision-making policies is a challenge in imitation learning (IL).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Towards VLA-Dreamer: Refining VLA Behavior Using World Models

The paper proposes a new architecture for Vision‑Language‑Action (VLA) models that improves sample efficiency by training a predictive world model on the vision encoder’s embedding space. It argues that these embeddings are action‑relevant and can be used to predict future states, addressing the lack of an explicit world model in current VLAs. The trained model can also support short‑term planning by sampling actions that lead to desired goal images.

By Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter