arXiv:2603. 19312v3 Announce Type: replace Abstract: Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse.
By Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero
arXiv:2605. 00412v3 Announce Type: replace Abstract: World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinforcement learning.
By Sen Cui, Jingheng Ma
arXiv:2608. 08553v1 Announce Type: cross Abstract: Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration.
By Rong Fu, Chunlei Meng, Yangchen Zeng, Xiaowen Ma, Yongtai Liu, Wangyu Wu, Shuo Yin, Zijian Zhang, Sicheng Li, Yingrui Ji, Chenhao Wang, Simon Fong
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence.
arXiv:2608.29029v3 Announce Type: replace-cross
Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free...
By Yanchen Huo, Ziying Song, Yadan Luo
CompAdapt is a physics-consistent text-to-video generation framework that extends diffusion-based models to handle composite physical behaviors such as coupled motions, multi-stage transitions, and multi-object collisions. It translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial parameters. The system introduces dynamics-aware prior matching for one-shot adaptation to new physical environments and a physics-aware latent feature fusion module to enhance visual fidelity during fast, complex motion, outperforming existing physics-constrained baselines on physics-focused T2V benchmarks.
By Haoran Qin (Harbin Institute of Technology, China), Renlong Wu (Harbin Institute of Technology, China), Tianyu Huang (Harbin Institute of Technology, China), Yukang Ding (Taobao, Alibaba Group, China), Hui Li (Harbin Institute of Technology, China), Wangmeng Zuo (Harbin Institute of Technology, China)
The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.
By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
arXiv:2603. 22282v2 Announce Type: replace-cross Abstract: We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
By Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu
The paper introduces a fusion‑aware hierarchical Gaussian patch representation that enables direct class‑guided generation of 3D Gaussian Splatting (3DGS) objects. By decomposing irregular Gaussian sets into canonical local patches and encoding them as structured tokens, the method fuses global class semantics with patch‑level geometry, appearance, spatial correspondence, and rendering‑sensitive cues. A structure‑aware rectified flow model, conditioned on patch positions and coupled with global‑local velocity prediction and density‑aware weighting, produces class‑conditioned 3DGS objects within seconds, achieving more coherent geometry, sharper local details, and better multi‑view consistency than baseline models.
By Yizhao Wang, Jingbo Wang, Guantao Zhang
Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck.
Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.
By Yanchen Huo, Ziying Song, Yadan Luo
arXiv:2606. 31050v1 Announce Type: cross Abstract: How to accurately predict a high-fidelity future world?
By Yujin Tang, Tian Zhou, Xin Lin, Cheng Tan, Yifan Hu, Rong Jin, SouYoung Jin, Liang Sun