Hugging Face Blog

Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains

arXiv Machine Learning
Sep 10

Mind the Gap: Navigating Inference with Optimal Transport Maps

The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.

By Malte Algren, Tobias Golling, Francesco Armando Di Bello, Christopher Pollard
arXiv Machine Learning
1d ago

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

The paper investigates how Mixture-of-Experts (MoE) Particle Transformers perform on the 188-class JetClass-II jet classification task. By varying expert count, routing capacity, top‑K, and auxiliary loss, the authors find that top‑1 MoE models can surpass dense baselines with similar nominal compute, but adding more experts yields diminishing accuracy gains. Activating multiple experts per token improves predictions at higher computational cost, and routing analyses reveal that expert assignments correlate with particle identity and kinematics, though this correlation does not consistently predict performance.

By Kaushik Pendiyala, Haris Zia, Trevin Lee, Timothy Legge, Alejandro J. De Leon, Zihan Zhao, Aaron Wang, Abhijith Gandrakota, Jennifer Ngadiuba, Richard Cavanaugh, Javier Duarte
arXiv Machine Learning
Jun 11

Robustness of Mixtures of Experts to Feature Noise

arXiv:2601. 14792v2 Announce Type: replace Abstract: Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling.

By Dong Sun, Rahul Nittala, Rebekka Burkholz
Hugging Face Trending Papers
Aug 24

Neural Boltzmann Equations

The dynamics of particles in the early universe are described by Boltzmann equations, which involve high-dimensional phase-space integrals. Classical approaches use quadrature integration and evolve t...

arXiv Machine Learning
Aug 27

Output Dilution: Redundant but Fragile Representations in MoE Models

The paper investigates how Mixture-of-Experts (MoE) models encode moral content compared to dense models. While linear probes can recover moral valence from almost all expert-layer combinations with high accuracy, these representations are far more fragile to activation noise, showing a 4.2‑fold drop in robustness. The authors attribute this fragility to output dilution: the MoE block averages across active experts, reducing the feedforward signal by nearly two orders of magnitude, which makes moral information vulnerable to perturbation even though routing remains stable.

By Orion Reblitz-Richardson