Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
Related stories
Mind the Gap: Navigating Inference with Optimal Transport Maps
The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.
Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers
The paper investigates how Mixture-of-Experts (MoE) Particle Transformers perform on the 188-class JetClass-II jet classification task. By varying expert count, routing capacity, top‑K, and auxiliary loss, the authors find that top‑1 MoE models can surpass dense baselines with similar nominal compute, but adding more experts yields diminishing accuracy gains. Activating multiple experts per token improves predictions at higher computational cost, and routing analyses reveal that expert assignments correlate with particle identity and kinematics, though this correlation does not consistently predict performance.
KIGNet: Physics-Motivated Multi-Graph Representation Learning for Explainable Jet Tagging
arXiv:2512. 07420v3 Announce Type: replace-cross Abstract: Jet identification plays a central role in analyzing data from high-energy collider experiments.
NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem
arXiv:2607. 19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology.
Simplex Demixing: Disentangling Multiple Light-Flavor Jets at Colliders
arXiv:2607. 24921v1 Announce Type: cross Abstract: Providing a practical and hadron-level definition of multiple jet flavors has been a long-standing challenge in collider physics.
Robustness of Mixtures of Experts to Feature Noise
arXiv:2601. 14792v2 Announce Type: replace Abstract: Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling.
An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning
arXiv:2607. 21702v1 Announce Type: new Abstract: Simulation-based inference (SBI) with machine learning is an increasingly important tool for solving inverse problems in science and engineering, including parameter inference and the inversion of detector effects.
Neural Boltzmann Equations
The dynamics of particles in the early universe are described by Boltzmann equations, which involve high-dimensional phase-space integrals. Classical approaches use quadrature integration and evolve t...
Pruning and Distilling Mixture-of-Experts into Dense Language Models
arXiv:2605. 28207v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment.
Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough
arXiv:2607. 10039v1 Announce Type: cross Abstract: Machine learning (ML) has become integral to fundamental physics, accelerating statistical workflows from data acquisition through inference and hypothesis testing.
Output Dilution: Redundant but Fragile Representations in MoE Models
The paper investigates how Mixture-of-Experts (MoE) models encode moral content compared to dense models. While linear probes can recover moral valence from almost all expert-layer combinations with high accuracy, these representations are far more fragile to activation noise, showing a 4.2‑fold drop in robustness. The authors attribute this fragility to output dilution: the MoE block averages across active experts, reducing the feedforward signal by nearly two orders of magnitude, which makes moral information vulnerable to perturbation even though routing remains stable.