arXiv Computer Vision By Wenhow Li (The Hong Kong University of Science and Technology), Chengwei MA (The Hong Kong University of Science and Technology), Hui Xiong (The Hong Kong University of Science and Technology), Ying-Cong Chen (The Hong Kong University of Science and Technology), Lei Zhang (The Hong Kong University of Science and Technology)

Platonic Representation Hypothesis on World Models

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

arXiv AI
Sep 1

Social-JEPA: Emergent Geometric Isomorphism

arXiv:2603.02263v3 Announce Type: replace-cross Abstract: World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such...

By Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou
arXiv Machine Learning
Jun 5

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

arXiv:2603. 19312v3 Announce Type: replace Abstract: Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse.

By Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero