arXiv:2609.06072v1 Announce Type: cross
Abstract: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-w...
By Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong
arXiv:2608. 07890v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert.
By Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao
ACE introduces a training‑free, calibration‑free framework for adaptive expert skipping in Mixture‑of‑Experts LLMs. It combines a Global Spectral Proxy that estimates global transformation capacity with a Router‑Conditioned Refinement that builds expert‑specific direction prototypes, enabling the model to skip low‑contribution experts while always keeping the top‑1 expert. Offline computation of expert statistics leaves only lightweight table lookups during inference, and experiments on three MoE‑based LLMs show ACE outperforms static and dynamic baselines, especially at high skipping ratios.
By Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li, Zhe Jiang, Dawei Yang
arXiv:2609.25655v1 Announce Type: new
Abstract: As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectu...
By Zhentao Tan, Chang Liu, Yao Liu, Yue Wu, Jieping Ye
arXiv:2510. 02345v4 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead.
By Peijun Zhu, Ning Yang, Baoliang Tian, Jiayu Wei, Weihao Zhang, Haijun Zhang, Pin Lv
MoRA is a framework for pruning Mixture-of-Experts (MoE) models by learning a router bias for each expert and optimizing it with a language‑modeling loss and a routing‑diversity regularizer. The learned biases sharpen routing distributions to identify critical experts and encourage diverse routing preferences. After pruning, MoRA uses an expert approximation mechanism that approximates the outputs of pruned experts with affine transformations of remaining experts, further improving performance.
By Yushuai Sun, Zikun Zhou, Lin Gao, Jun Yu, Wenjie Pei