Mixture of Experts Explained
Related stories
Mixture of Experts (MoEs) in Transformers
Machine Learning Experts - Sasha Luccioni
Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
Machine Learning Experts - Margaret Mitchell
Robustness of Mixtures of Experts to Feature Noise
arXiv:2601. 14792v2 Announce Type: replace Abstract: Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling.
PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization
arXiv:2606. 03444v1 Announce Type: cross Abstract: Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent in monolithic distillation.
Machine Learning Experts - Lewis Tunstall
Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
arXiv:2607. 20426v1 Announce Type: cross Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal knowledge or have poor cross-domain generalization.
A theoretical model for task routing in mixture-of-expert transformers
arXiv:2606. 14398v1 Announce Type: new Abstract: Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed.
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
arXiv:2608. 15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure.
Does Role Specialization Matter for Explanation Faithfulness in Mixture-of-Experts?
arXiv:2606. 29613v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have recently been extended with role-based mechanisms for interpretability.