arXiv AI By Minghao Yang, Ren Togo, Guang Li, Takahiro Ogawa, Miki Haseyama

L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts

Read the original on arXiv AI →

The paper introduces L2R, a routing framework for Mixture-of-Experts models that reshapes the routing space into a shared low‑rank latent space and employs Saturated Inner‑Product Scoring to control Lipschitz behavior, resulting in smoother and more stable routing geometry. It also adds a parameter‑efficient multi‑anchor routing mechanism to increase expert expressiveness. Experiments on an OLMoE‑based language model and a ViT‑based ImageNet setting demonstrate improved overall performance and better routing geometry and expert discrimination.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv Machine Learning
1d ago

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

MoRA is a framework for pruning Mixture-of-Experts (MoE) models by learning a router bias for each expert and optimizing it with a language‑modeling loss and a routing‑diversity regularizer. The learned biases sharpen routing distributions to identify critical experts and encourage diverse routing preferences. After pruning, MoRA uses an expert approximation mechanism that approximates the outputs of pruned experts with affine transformations of remaining experts, further improving performance.

By Yushuai Sun, Zikun Zhou, Lin Gao, Jun Yu, Wenjie Pei