arXiv:2608. 16245v1 Announce Type: new Abstract: Disentangled representation learning seeks latent representations whose indicidual dimensions each align with a distinct covariate.
By Ma{\l}gorzata {\L}az\k{e}cka, Ewa Szczurek
The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.
By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv:2606. 15306v1 Announce Type: cross Abstract: We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared across those tasks and use it to improve future decisions.
By Daksh Mittal, Tommaso Castellani, Thomson Yen, Naimeng Ye, Fangyu Wu, Minghui Chen, Tiffany Cai, Emmanouil Koukoumidis, William Zeng, Hongseok Namkoong
arXiv:2401. 04890v2 Announce Type: replace-cross Abstract: This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interest depend sparsely on observed auxiliary variables and/or past latent factors.
By S\'ebastien Lachapelle, Pau Rodr\'iguez L\'opez, Yash Sharma, Katie Everett, R\'emi Le Priol, Alexandre Lacoste, Simon Lacoste-Julien
Joint-embedding predictive architectures (JEPAs) learn latent dynamics for planning and avoid representation collapse by matching features to maximum-entropy distributions such as isotropic Gaussians,...
arXiv:2608. 11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informative for parameter calibration.
By Qiyao Zhou, Xujia Zhu, Pierre Joli, Yu Cong, Sibo Cheng
The paper introduces a lightweight Fourier auxiliary head to enforce physically-informed structuring of latent states in JEPA-style world models, addressing a newly identified failure mode called physical representation laziness that hampers planning in dynamic environments. Experiments show that this auxiliary supervision improves planning success rates, enhances latent space correlations with key physical properties, and boosts data efficiency, even when the baseline model does not exhibit laziness.
By Penghao Zhu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv:2607. 04409v1 Announce Type: new Abstract: Learning and planning in imagination using world models provides an effective paradigm for training agents for decision-making.
By Fan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen, Guangyi Chen, Kevin Murphy, Biwei Huang, Kun Zhang
arXiv:2608. 20065v1 Announce Type: new Abstract: World models construct latent states that support prediction, planning, and reasoning about an underlying system.
By Taoyong Cui, Pheng Ann Heng, Wanli Ouyang
arXiv:2607. 25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks.
By Xiaoyu Huang, Lulu Wang
arXiv:2603.20111v2 Announce Type: replace
Abstract: The Joint-Embedding Predictive Architecture (JEPA) is often seen as a non-generative alternative to likelihood-based self-supervised learning, emph...
By Moritz G\"ogl, Christopher Yau
HALO introduces a hyperspherical VAE to constrain continuous latent representations to a fixed‑radius shell, stabilizing numerical fluctuations. It then employs a masked autoregressive model that balances parallel decoding with temporal correlation learning, reducing inference steps and improving stability. Experiments show HALO achieves state‑of‑the‑art generation performance with significantly better inference efficiency compared to existing baselines.
By Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang