arXiv AI By Ke Fang, Yupu Yao, Lu Cheng

ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning

Read the original on arXiv AI →

The paper introduces ATLAS, a training objective that preserves relational geometry in latent world models while calibrating the global latent distribution. By transferring normalized pairwise structure from an informative encoder to the planning latent and applying Wasserstein embedding matching, ATLAS improves goal‑reaching success on tasks such as PushT, TwoRoom, and OGBench‑Cube, especially on higher‑novelty episodes. Diagnostics show stronger novelty‑related structure, better marginal calibration, and lower multi‑step prediction error in the planning latent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to improve goal‑conditioned robotic planning. By preventing latent collapse and grounding representations in physical configuration, the model achieves top success rates on tasks such as TwoRoom, PushT, and OGBench‑Cube, outperforming the baseline LeWorldModel. Ablation studies confirm that state alignment consistently boosts planning success over inverse dynamics alone across all four benchmark tasks.

By Muyuan Liu (GENISOM AI, Beijing, China), Yue Huang (GENISOM AI, Beijing, China), Zheng Liang (GENISOM AI, Beijing, China), Xiang Gao (GENISOM AI, Beijing, China)
arXiv Computer Vision
3d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin