The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.
By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv:2606. 07414v1 Announce Type: new Abstract: Sparsity allows scaling model parameters without proportionally increasing computational cost.
By Simon Schug
arXiv:2609.37533v1 Announce Type: new
Abstract: Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically...
By Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov
arXiv:2602. 17554v3 Announce Type: replace Abstract: Training large-scale generative models is resource-intensive and relies heavily on heuristic dataset weighting.
By Corinna Cortes, Mehryar Mohri, Yutao Zhong
arXiv:2609.09241v1 Announce Type: cross
Abstract: Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large f...
By Dohyeon Kim, Bedionita Soro, Sung Ju Hwang
arXiv:2606. 09885v1 Announce Type: new Abstract: Mixture-of-Experts large language models (LLMs) scale efficiently through sparse activation, yet their deployment is fundamentally constrained by the large static parameter footprint of experts.
By Jiangyang He, Shaolin Zhu, Deyi Xiong
arXiv:2609.21672v1 Announce Type: new
Abstract: Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeab...
By Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen
arXiv:2606. 01062v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge.
By Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu, Yinglong Xia, Qiang Zhang, Qifan Wang, Ren Chen, Dongqi Fu, Jiayi Liu, Zhoukai Zhao, Xiangjun Fan, Benyu Zhang, Yixin Chen
The paper investigates how Mixture-of-Experts (MoE) models encode moral content compared to dense models. While linear probes can recover moral valence from almost all expert-layer combinations with high accuracy, these representations are far more fragile to activation noise, showing a 4.2‑fold drop in robustness. The authors attribute this fragility to output dilution: the MoE block averages across active experts, reducing the feedforward signal by nearly two orders of magnitude, which makes moral information vulnerable to perturbation even though routing remains stable.
By Orion Reblitz-Richardson
arXiv:2606. 30355v1 Announce Type: cross Abstract: As real-world prediction systems often face missing modalities at inference, incomplete multimodal learning (IML) remains a practical challenge.
By Seunghun Baek, Jihwan Park, Jaeyoon Sim, Minjae Jeong, Hoseok Lee, Won Hwa Kim
arXiv:2606. 20544v1 Announce Type: new Abstract: Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities.
By Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa, Anqi Liu
arXiv:2606. 17952v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) architectures enable scaling LLM parameters under a fixed inference budget by activating only a small subset of experts via top-$k$ routing.
By Miko{\l}aj Zasada, {\L}ukasz Struski, Jacek Tabor, Marcin Kurdziel