arXiv Machine Learning By Kaushik Pendiyala, Haris Zia, Trevin Lee, Timothy Legge, Alejandro J. De Leon, Zihan Zhao, Aaron Wang, Abhijith Gandrakota, Jennifer Ngadiuba, Richard Cavanaugh, Javier Duarte

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

Read the original on arXiv Machine Learning →

The paper investigates how Mixture-of-Experts (MoE) Particle Transformers perform on the 188-class JetClass-II jet classification task. By varying expert count, routing capacity, top‑K, and auxiliary loss, the authors find that top‑1 MoE models can surpass dense baselines with similar nominal compute, but adding more experts yields diminishing accuracy gains. Activating multiple experts per token improves predictions at higher computational cost, and routing analyses reveal that expert assignments correlate with particle identity and kinematics, though this correlation does not consistently predict performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation

The paper studies how the design of Mixture-of-Experts (MoE) routers affects inference speed when combined with Speculative Decoding (SD). It shows that routers promoting high expert coactivation reduce memory transfer costs and improve runtime. By integrating a global load‑balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism, the authors achieve a 21% throughput gain over baseline MoEs while preserving accuracy.

By Kumari Nishu, Han-Byul Kim, Santosh Chilkunda, Maxwell Horton, Arnav Kundu, Mohammad Samragh, Lauren Hannah, Mohammad Sekhavat, Nikhil Bhendawade, Manuel Ciosici, Iman Mirzadeh, Keivan Alizadeh Vahid, David Harrison, Irina Belousova, Mehrdad Farajtabar, Minsik Cho
arXiv Machine Learning
Jul 24

PreMoE: Proactive Inference for Efficient Mixture-of-Experts

arXiv:2505. 17639v4 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization.

By Zehua Pei, Ying Zhang, Hui-Ling Zhen, Tao Yuan, Xianzhi Yu, Zhenhua Dong, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu