arXiv Machine Learning By Leonardo F. Toso, Davit Shadunts, Yunyang Lu, Gloria Geng, Nihal Sharma, Donglin Zhan, Nam H. Nguyen, James Anderson

JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models

Read the original on arXiv Machine Learning →

JEPA-Bisim introduces a bisimulation encoder to joint-embedding predictive world models, ensuring that states with similar transition dynamics are mapped to nearby latent representations while suppressing irrelevant slow features such as background changes and distractors. The approach improves robustness on navigation (PointMaze) and manipulation (PushT) tasks under varied test-time visual conditions, achieving up to tenfold smaller latent spaces than DINO-WM. It remains effective across different pre-trained visual encoders, including DINOv2, SimDINOv2, and iBOT.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

DeepJEPA: Scaling World Models from Within

DeepJEPA is a weight‑tied joint‑embedding predictive world model that treats transition depth as an inner test‑time scaling axis, learning when additional recurrent updates are worthwhile for each candidate and rollout step. Unlike traditional planners that uniformly deepen every transition, DeepJEPA concentrates extra computation on decision‑critical events such as contact onset and sustained object interaction, achieving comparable or better performance with only 1.00–1.26 updates per transition across five visual‑control settings. The approach demonstrates that improved planning does not require uniformly better object‑state decodability, but rather targeted internal computation where it can alter the planner’s elite set and action selection.

By Zijian Jin, Yunbei Zhang, Yuanzhe Liu, Ming Liu, Baian Chen, Weirui Ye, Shilong Liu, Marco Pavone
arXiv Computer Vision
6d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin