arXiv Machine Learning

Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts

arXiv:2606. 19036v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networks.

arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv AI
Sep 3

Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts

The paper investigates how routing decisions in sparse mixture-of-experts (MoE) models evolve across layers. By aligning router control subspaces with generalized orthogonal Procrustes analysis, the authors find that a single linear transition can predict routing states across depth with substantial accuracy, revealing a shared geometric structure. They further demonstrate that these canonical states preserve expert selection better than generic hidden representations and improve next‑step routing predictions, reducing negative log‑likelihood by up to 15.7% on OLMoE and 6.2% on Phi.

By Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov
arXiv Machine Learning
Sep 21

Schedule optimization for tau-leaping in masked discrete diffusion

The paper studies how to choose sampling schedules for tau‑leaping in masked discrete diffusion models. By deriving an exact integral representation of the factorization error ε_fact in terms of a dependence density ρ, the authors develop estimators and recursive equations that identify the unique optimal schedule under a monotonicity condition. In the large‑scale limit, they provide explicit characterizations of the optimal smooth schedule and show that while optimizing smooth schedules can improve constants, it does not change the N/K scaling unless the dependence density degenerates, in which case asymptotic improvements are possible.

By Cecilia Secchi, Giacomo Zanella
arXiv Machine Learning
Jul 14

Sharp Concentration Bounds for Bundle-Valued Statistics on Manifolds

arXiv:2607. 10592v1 Announce Type: new Abstract: Many geometric statistics and manifold learning pipelines routinely produce observations -- such as tangent vectors or local frames -- whose natural home is a varying family of fibers attached to different points of a base manifold, rather than a single shared vector space.

By Swagatam Das, Vaclav Snasel
arXiv Statistics ML
3d ago

The Geometry of Randomized Smoothing on Feasible Sets

arXiv:2609.39497v1 Announce Type: cross Abstract: Randomized smoothing certifies the probability of a fixed output event as the center of Gaussian noise moves. Feasibility or confidence filtering rep...

By Syed Izhan Khilji, Alireza Furutanpey, Schahram Dustdar
arXiv Machine Learning
Sep 22

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng