arXiv AI By Petr Ivashkov, Randall Balestriero, Bernhard Sch\"olkopf

Sensorimotor World Models: Perception for Action via Inverse Dynamics

Read the original on arXiv AI →

arXiv:2606. 20104v1 Announce Type: cross Abstract: Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 1

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.

By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
Hugging Face Trending Papers
Sep 3

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to better support goal‑conditioned robotic planning. By incorporating inverse dynamics, the model prevents latent collapse and encodes action information, while state alignment ties consecutive latent states to their physical configurations and motions. Experiments on four benchmark tasks show the model achieves top success rates on TwoRoom, PushT, and OGBench‑Cube, and its state alignment consistently improves planning performance over inverse dynamics alone.