The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.
By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv:2605. 17232v5 Announce Type: replace Abstract: Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology.
By Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Markos A. Katsoulakis
arXiv:2607. 20540v1 Announce Type: cross Abstract: How should a diffusion model decide which noise levels to train on, and how much?
By Luca Ambrogioni, Giulio Franzese, Alberto Foresti, Gabriel Raya, Bac Nguyen, Georgios Batzolis, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai, Yuki Mitsufuji
The paper investigates how routing decisions in sparse mixture-of-experts (MoE) models evolve across layers. By aligning router control subspaces with generalized orthogonal Procrustes analysis, the authors find that a single linear transition can predict routing states across depth with substantial accuracy, revealing a shared geometric structure. They further demonstrate that these canonical states preserve expert selection better than generic hidden representations and improve next‑step routing predictions, reducing negative log‑likelihood by up to 15.7% on OLMoE and 6.2% on Phi.
By Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov
The paper studies how to choose sampling schedules for tau‑leaping in masked discrete diffusion models. By deriving an exact integral representation of the factorization error ε_fact in terms of a dependence density ρ, the authors develop estimators and recursive equations that identify the unique optimal schedule under a monotonicity condition. In the large‑scale limit, they provide explicit characterizations of the optimal smooth schedule and show that while optimizing smooth schedules can improve constants, it does not change the N/K scaling unless the dependence density degenerates, in which case asymptotic improvements are possible.
By Cecilia Secchi, Giacomo Zanella
arXiv:2505.20817v3 Announce Type: replace-cross
Abstract: Gradient clipping is widely used in language-model training to control heavy-tailed gradient noise and can improve convergence guarantees ove...
By Taha El Bakkali El Kadi, Savelii Chezhegov, Aleksandr Beznosikov, Samuel Horv\'ath, Eduard Gorbunov