arXiv Computer Vision
Aug 26

Platonic Representation Hypothesis on World Models

arXiv:2608.23720v1 Announce Type: new Abstract: World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundamen...

By Wenhow Li (The Hong Kong University of Science and Technology), Chengwei MA (The Hong Kong University of Science and Technology), Hui Xiong (The Hong Kong University of Science and Technology), Ying-Cong Chen (The Hong Kong University of Science and Technology), Lei Zhang (The Hong Kong University of Science and Technology)
arXiv AI
Jun 17

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

arXiv:2606. 17536v1 Announce Type: cross Abstract: Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry.

By Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang
Hugging Face Trending Papers
Jul 9

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors.