ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
arXiv:2603. 22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.
arXiv:2603. 22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.
WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.
arXiv:2609.21379v1 Announce Type: new Abstract: Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video ge...
arXiv:2512.21004v2 Announce Type: replace Abstract: Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregress...
arXiv:2606. 14765v1 Announce Type: cross Abstract: Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning.
arXiv:2608.24855v1 Announce Type: new Abstract: Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online traje...
AWM‑VLA introduces a unified framework that embeds aligned world modeling directly into a diffusion‑transformer vision‑language‑action policy. By adding learnable future tokens aligned with vision‑language embeddings of future observations, the policy can anticipate long‑term consequences while generating actions. The method extends this with an object‑centric alignment objective and a principled weighting scheme, achieving up to 21% higher success rates on RoboCasa and humanoid tabletop benchmarks and producing object‑centric rationales preferred by human raters in 83% of cases.
arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.
arXiv:2606. 05758v1 Announce Type: cross Abstract: Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens.
arXiv:2607. 27924v1 Announce Type: new Abstract: In the physical world we inhabit, space and time are fundamentally continuous.
arXiv:2608. 15869v1 Announce Type: cross Abstract: Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments.