Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
Related stories
KIGNet: Physics-Motivated Multi-Graph Representation Learning for Explainable Jet Tagging
arXiv:2512. 07420v3 Announce Type: replace-cross Abstract: Jet identification plays a central role in analyzing data from high-energy collider experiments.
NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem
arXiv:2607. 19406v1 Announce Type: new Abstract: Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology.
Simplex Demixing: Disentangling Multiple Light-Flavor Jets at Colliders
arXiv:2607. 24921v1 Announce Type: cross Abstract: Providing a practical and hadron-level definition of multiple jet flavors has been a long-standing challenge in collider physics.
Robustness of Mixtures of Experts to Feature Noise
arXiv:2601. 14792v2 Announce Type: replace Abstract: Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling.
An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning
arXiv:2607. 21702v1 Announce Type: new Abstract: Simulation-based inference (SBI) with machine learning is an increasingly important tool for solving inverse problems in science and engineering, including parameter inference and the inversion of detector effects.
Pruning and Distilling Mixture-of-Experts into Dense Language Models
arXiv:2605. 28207v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment.
Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough
arXiv:2607. 10039v1 Announce Type: cross Abstract: Machine learning (ML) has become integral to fundamental physics, accelerating statistical workflows from data acquisition through inference and hypothesis testing.
Stable Global Weighting of Flow Mixtures using Simplex Exponential Moving Average
arXiv:2607. 03809v1 Announce Type: new Abstract: Normalising flows provide a powerful variational family for approximate inference, yet individual architectures often fail to generalise across heterogeneous posterior geometries.
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
arXiv:2608. 03457v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood.
Leveraging generative models to assist Monte Carlo sampling
arXiv:2608. 07648v1 Announce Type: cross Abstract: Sampling high-dimensional probability distributions is a central task in scientific computing, with applications ranging from Bayesian inference to statistical physics and molecular simulation.