Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing
arXiv:2609. 17940v1 Announce Type: new Abstract: Sparse mixture-of-experts models route each token through a sequence of expert selections.
The paper investigates how Mixture-of-Experts (MoE) models encode moral content compared to dense models. While linear probes can recover moral valence from almost all expert-layer combinations with high accuracy, these representations are far more fragile to activation noise, showing a 4.2‑fold drop in robustness. The authors attribute this fragility to output dilution: the MoE block averages across active experts, reducing the feedforward signal by nearly two orders of magnitude, which makes moral information vulnerable to perturbation even though routing remains stable.
arXiv:2609. 17940v1 Announce Type: new Abstract: Sparse mixture-of-experts models route each token through a sequence of expert selections.
arXiv:2604.23036v2 Announce Type: replace-cross Abstract: Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layer...
arXiv:2609.38823v1 Announce Type: new Abstract: Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly becau...
The paper studies how the design of Mixture-of-Experts (MoE) routers affects inference speed when combined with Speculative Decoding (SD). It shows that routers promoting high expert coactivation reduce memory transfer costs and improve runtime. By integrating a global load‑balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism, the authors achieve a 21% throughput gain over baseline MoEs while preserving accuracy.
arXiv:2606. 27866v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models.
The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.
arXiv:2607. 28097v1 Announce Type: new Abstract: Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions.
arXiv:2609.36222v1 Announce Type: new Abstract: Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferrin...
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.
MetaNet is a support‑set controller that predicts, for each layer of a Mixture‑of‑Experts model, an expert‑retention threshold and a bounded routing bias while keeping the backbone, experts, and router frozen. On DeepSeek‑MoE‑16B‑Chat, MetaNet offers a tunable trade‑off between accuracy and expert activation: a conservative setting activates 3.61 experts on average (40% fewer than a fixed k=6) with comparable MMLU accuracy, whereas an aggressive setting activates only 2.28 experts (62% fewer) with a modest accuracy drop. The MMLU‑trained controller also transfers to C‑Eval, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.
arXiv:2607. 24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory.
arXiv:2606. 10338v1 Announce Type: cross Abstract: Machine unlearning is increasingly important for large language models, yet unlearning in Mixture-of-Experts (MoE) architectures remains underexplored.