arXiv:2607. 28308v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions.
By Huiyuan Tian, Bonan Xu, Shijian Li
arXiv:2608. 07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices.
By Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao
IntBMoE introduces a block‑conditioned mixture‑of‑experts that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. Each internal layer uses a lightweight hypernetwork to merge all expert bases into a single composed expert, while a router selects only a few blocks per token, keeping compute and memory costs low. Experiments on image classification, language modeling, and sequential recommendation demonstrate consistent performance gains, and the model is deployed in AMap’s generative recommendation system, improving UVCTR by 2.4% in online A/B tests.
By Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu
arXiv:2608. 08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs.
By Zongfei Li
The paper introduces two token-error supervision methods for sparse mixture-of-experts (MoE) large language models, aligning routing affinities with token-level cross-entropy loss. The first method predicts an error score per expert and uses it to adjust affinities before top‑K selection, while the second directly aligns router affinities to the model’s objective without extra heads or inference changes. Experiments on Granite and ARC‑Challenge show accuracy gains of about 2.3–2.94 percentage points over a parameter‑matched baseline, preserving the native sparse execution budget.
By Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit
The paper examines how expert pruning—removing low‑importance experts in Mixture‑of‑Experts models—fails when the router is over‑dispersed, a condition caused by aggressive load‑balancing that spreads tokens nearly uniformly across experts. In this regime, traditional importance signals from router probabilities collapse, making perplexity an unreliable predictor of downstream accuracy; for example, the lowest‑perplexity pruning on gpt‑oss‑20B harms mathematical reasoning while the highest‑perplexity pruning preserves it. To address this, the authors introduce Minimax Expert Score Allocation (MESA), a domain‑aware method that iteratively boosts scores for the most affected domain, achieving minimal worst‑case degradation across domains and outperforming baseline pruning strategies on multiple benchmarks while reducing memory usage.
By Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade