arXiv Machine Learning By Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen, Nhat-Tri Ho, Viet-Hoang Tran, Tho Quan, Tan Minh Nguyen

Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts

Read the original on arXiv Machine Learning →

arXiv:2606. 19036v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv AI
Sep 3

Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts

The paper investigates how routing decisions in sparse mixture-of-experts (MoE) models evolve across layers. By aligning router control subspaces with generalized orthogonal Procrustes analysis, the authors find that a single linear transition can predict routing states across depth with substantial accuracy, revealing a shared geometric structure. They further demonstrate that these canonical states preserve expert selection better than generic hidden representations and improve next‑step routing predictions, reducing negative log‑likelihood by up to 15.7% on OLMoE and 6.2% on Phi.

By Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov
arXiv Machine Learning
Sep 21

Schedule optimization for tau-leaping in masked discrete diffusion

The paper studies how to choose sampling schedules for tau‑leaping in masked discrete diffusion models. By deriving an exact integral representation of the factorization error ε_fact in terms of a dependence density ρ, the authors develop estimators and recursive equations that identify the unique optimal schedule under a monotonicity condition. In the large‑scale limit, they provide explicit characterizations of the optimal smooth schedule and show that while optimizing smooth schedules can improve constants, it does not change the N/K scaling unless the dependence density degenerates, in which case asymptotic improvements are possible.

By Cecilia Secchi, Giacomo Zanella