arXiv:2608.20974v1 Announce Type: cross
Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent featu...
By Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
Drive‑HWM introduces a hierarchical slow‑fast world modeling framework for autonomous driving. The slow model predicts multi‑step future representations, while the fast model jointly predicts the next frame and immediate action using a lightweight multimodal backbone and an autoregressive expert. Dynamic‑Aware Latents, learned through optical‑flow prediction, explicitly capture motion dynamics, and experiments on NAVSIM v1 and v2 show strong driving performance with validated ablation studies.
By Zhaoxin Fan, Tianbao Zhang, Wenjun Wu, Xiaofeng Wang, Yeying Jin, Jian Zhao, Zheng Zhu, Shuicheng Yan
arXiv:2608. 03244v1 Announce Type: new Abstract: Image-goal visual navigation is a fundamental capability for embodied agents.
By Changqing Zhou, Yueru Luo, Zeyu Jiang, Changhao Chen
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
By Chen Yang, Yuhao Wei, Ze Xu, Ziheng Zou, Shuang Liang, Delin Ouyang, Lingfeng Qi, Jie Li, Guofa Li
PhysWAM is a unified world-action model for autonomous driving that jointly denoises multiview video, metric depth, and ego motion using a flow‑matching transformer. It introduces Coupled Point Projection (CPP), a geometric constraint that aligns generated depth points with LiDAR data after applying the predicted SE(3) ego motion, thereby enforcing physical consistency. At inference, trajectory selection uses a simple label‑free consensus rule, and the model demonstrates strong planning performance, zero‑shot transfer to unseen environments, and accurate, temporally coherent depth and video predictions.
By Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang
DA‑WAM is a framework that integrates predictive representation learning, action‑conditioned future modeling, and trajectory scoring into a single decision‑making objective for autonomous driving. It uses an online encoder with a stable momentum target to keep future representations aligned with the driving task, generating a distinct future latent for each trajectory candidate. A future‑latent‑conditioned scorer evaluates these latents, with expert‑matched trajectories supervised by observed futures and safety‑critical hard negatives providing additional guidance, achieving state‑of‑the‑art results on NAVSIM‑v1 and NAVSIM‑v2.
By Ruiguo Zhong, Benshan Ma, Xiaolong Chen, Lang Zhang, Mingyue Feng, Yaonong Wang, Pei Liu, Jun Ma