Mixtral of experts
Related stories
Machine Learning Experts - Sasha Luccioni
Mixture of Experts (MoEs) in Transformers
Machine Learning Experts - Margaret Mitchell
Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization
arXiv:2606. 03444v1 Announce Type: cross Abstract: Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent in monolithic distillation.
Reliable Fusion of Conflicting Experts
The paper introduces a probabilistic‑circuit framework for fusing opinions from multiple black‑box experts in noisy, conflict‑prone environments. It dynamically assigns context‑specific credibility to each expert, allowing reliable aggregation without needing access to their internal models or retraining. Experiments on multiple‑choice question answering with large language models show that this method outperforms individual models and static ensemble baselines, consistently improving predictive accuracy and decision reliability under disagreement.
Decomposing LLM-Judge Uncertainty to Target Expert Labels
arXiv:2609.06444v2 Announce Type: cross Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties:...
Machine Learning Experts - Lewis Tunstall
ResMerge: Residual-based Spectral Merging of Large Language Models
ResMerge is a new framework for merging large language models trained via reinforcement learning. It separates each model’s task vector into a leading spectral head and a residual component, finding that both parts contain valuable behavior knowledge but behave differently during merging. The method builds a stable residual backbone using Spherical Residual Consensus Adaptation and then adds a lightweight head correction module that activates only when experts agree, leading to better preservation of expert capabilities compared to existing merging baselines.
Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts
In current Mixture-of-Experts architectures, routing is performed based on representations dominated by structure shared across all tokens, limiting expert specialization. We show that contrasting eac...
Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA
arXiv:2607. 26052v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$.