Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.
The paper investigates how Mixture-of-Experts (MoE) Particle Transformers perform on the 188-class JetClass-II jet classification task. By varying expert count, routing capacity, top‑K, and auxiliary loss, the authors find that top‑1 MoE models can surpass dense baselines with similar nominal compute, but adding more experts yields diminishing accuracy gains. Activating multiple experts per token improves predictions at higher computational cost, and routing analyses reveal that expert assignments correlate with particle identity and kinematics, though this correlation does not consistently predict performance.
arXiv:2512. 07420v3 Announce Type: replace-cross Abstract: Jet identification plays a central role in analyzing data from high-energy collider experiments.
arXiv:2607. 19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology.
arXiv:2607. 24921v1 Announce Type: cross Abstract: Providing a practical and hadron-level definition of multiple jet flavors has been a long-standing challenge in collider physics.