Hugging Face Trending Papers

Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?

Read the original on Hugging Face Trending Papers →

S-JEPA uses soft Gaussian mixture model (GMM) posteriors to preserve uncertainty in encoder representations. The study compares three mapping strategies—real soft, fixed-random permutation, and uniform tail—to determine whether the assignment of non‑maximal probabilities to GMM components influences learned representations. Results show that the real soft mapping outperforms the controls, indicating that the numerical probability structure alone does not fully determine the encoder’s learned representation, and that the mapping of non‑maximal probabilities to GMM components matters.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Aug 20

Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?

S-JEPA employs soft Gaussian mixture model (GMM) posteriors to preserve uncertainty in encoder representations. The study compares three mapping strategies—real soft, fixed-random permutation, and uniform tail—to assess whether the assignment of non‑maximal probabilities to GMM components influences learned representations. Results show that the real soft mapping outperforms the controls, indicating that the numerical probability structure alone does not fully determine the encoder’s learned representation.

By Wenxuan He, Yunpeng Li, Shan Liang
Hugging Face Trending Papers
Aug 18

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.

arXiv Machine Learning
Jun 9

Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps

arXiv:2510. 12744v2 Announce Type: replace-cross Abstract: We develop a unified statistical framework for softmax-gated Gaussian mixture of experts (SGMoE) that addresses three long-standing obstacles in parameter estimation and model selection: (i) non-identifiability of gating parameters up to common translations, (ii) intrinsic gate-expert interactions that induce coupled differential relations in the likelihood, and (iii) the tight numerator-denominator coupling in the softmax-induced conditional density.

By Do Tien Hai, Trung Nguyen Mai, TrungTin Nguyen, Nhat Ho, Binh T. Nguyen, Christopher Drovandi