arXiv Machine Learning By Wenxuan He, Yunpeng Li, Shan Liang

Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?

Read the original on arXiv Machine Learning →

S-JEPA employs soft Gaussian mixture model (GMM) posteriors to preserve uncertainty in encoder representations. The study compares three mapping strategies—real soft, fixed-random permutation, and uniform tail—to assess whether the assignment of non‑maximal probabilities to GMM components influences learned representations. Results show that the real soft mapping outperforms the controls, indicating that the numerical probability structure alone does not fully determine the encoder’s learned representation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 19

Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?

S-JEPA uses soft Gaussian mixture model (GMM) posteriors to preserve uncertainty in encoder representations. The study compares three mapping strategies—real soft, fixed-random permutation, and uniform tail—to determine whether the assignment of non‑maximal probabilities to GMM components influences learned representations. Results show that the real soft mapping outperforms the controls, indicating that the numerical probability structure alone does not fully determine the encoder’s learned representation, and that the mapping of non‑maximal probabilities to GMM components matters.

Hugging Face Trending Papers
Aug 18

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.

arXiv Machine Learning
Jun 5

Equivariant Neural Belief Propagation

arXiv:2606. 06344v1 Announce Type: new Abstract: Probabilistic inference over spatially embedded variables requires beliefs that respect $SE(3)$ symmetry, yet existing equivariant networks produce only scalars and vectors -- not the rank-2 precision tensors needed for anisotropic uncertainty, and single-component messages collapse multi-modal energy landscapes to physically meaningless averages.

By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv AI
Aug 19

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that replaces the Gaussian regularizer used in Joint‑Embedding Predictive Architectures (JEPAs) with a contrastive inverse‑dynamics head. AC‑MTM trains a forward latent‑prediction model while an auxiliary inverse‑dynamics task forces the encoder to distinguish actions from latent transitions, preventing collapse without requiring a target network or reconstruction loss. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM matches or surpasses the performance of the Gaussian‑based SIGReg regularizer, achieving up to 20–24 point improvements on the OGBench Visual Scene benchmark.

By Jack Boylan, Chris Hokamp
arXiv Machine Learning
Jun 9

Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps

arXiv:2510. 12744v2 Announce Type: replace-cross Abstract: We develop a unified statistical framework for softmax-gated Gaussian mixture of experts (SGMoE) that addresses three long-standing obstacles in parameter estimation and model selection: (i) non-identifiability of gating parameters up to common translations, (ii) intrinsic gate-expert interactions that induce coupled differential relations in the likelihood, and (iii) the tight numerator-denominator coupling in the softmax-induced conditional density.

By Do Tien Hai, Trung Nguyen Mai, TrungTin Nguyen, Nhat Ho, Binh T. Nguyen, Christopher Drovandi