Environment-Robust Representation Learning with Empirical Bayes
arXiv:2606. 05365v1 Announce Type: cross Abstract: We consider multi-environment prediction problems.
arXiv:2506. 22675v4 Announce Type: replace-cross Abstract: Invariant prediction [Peters et al.
arXiv:2606. 05365v1 Announce Type: cross Abstract: We consider multi-environment prediction problems.
arXiv:2608.29029v3 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free...
arXiv:2607. 18209v1 Announce Type: cross Abstract: This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments.
arXiv:2601. 02322v2 Announce Type: replace-cross Abstract: A common approach to out-of-distribution prediction restricts models to causal or invariant covariates to avoid spurious associations that may change across environments.
This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments. The joint distribution of the covariates may vary across environments, whereas the latent structure is decomposed into invariant factors with shared loadings and heterogeneous factors with environment-specific loadings.
arXiv:2606. 15458v1 Announce Type: cross Abstract: Variational inference (VI) is a core engine of modern AI, enabling scalable approximate Bayesian learning and uncertainty-aware training of large probabilistic and generative models.
The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.
arXiv:2410. 14843v4 Announce Type: replace-cross Abstract: Vanilla variational inference finds an optimal approximation to the Bayesian posterior distribution, but even the exact Bayesian posterior is often not meaningful under model misspecification.
arXiv:2603.20111v2 Announce Type: replace Abstract: The Joint-Embedding Predictive Architecture (JEPA) is often seen as a non-generative alternative to likelihood-based self-supervised learning, emph...
Flow-JEPA introduces a conditional flow matching dynamics model that generates a sequence of future latent states conditioned on current observations and actions, replacing deterministic autoregressive prediction with stochastic trajectory-level prediction. By using a Gaussian flow source, the model learns to transport perturbed latent trajectories toward clean future representations while remaining within the reconstruction‑free JEPA framework. The approach improves mean success rates from 86% to 92% under clean observations and from 67% to 86% under noisy conditions.
The paper presents PAC‑Bayesian reconstruction guarantees for Variational Autoencoders applied to time‑series data. It extends existing bounds, which were limited to i.i.d. settings, to Markovian latent structures, allowing temporal dependencies to be captured without the bounds growing with trajectory length. The authors also provide an example framework showing that the required assumptions are not overly restrictive.
The paper introduces prequential posteriors, a Bayesian approach that uses a predictive‑sequential loss function to update deep generative forecasting models (DGFMs) when new data arrive. By adopting a consistency notion suitable for model misspecification, the authors prove that both the loss minimizer and the posterior concentrate on parameters with optimal predictive performance. Scalable inference is achieved with parallelisable waste‑free sequential Monte Carlo samplers that employ preconditioned gradient kernels, and the method is validated on synthetic and real meteorological time‑series data.