arXiv Machine Learning By Yongli Xiang, Vinoth Nandakumar, Yunzhi Yao, Peike Li, Tongliang Liu

A theoretical model for task routing in mixture-of-expert transformers

Read the original on arXiv Machine Learning →

arXiv:2606. 14398v1 Announce Type: new Abstract: Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 2

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

arXiv:2606. 01062v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge.

By Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu, Yinglong Xia, Qiang Zhang, Qifan Wang, Ren Chen, Dongqi Fu, Jiayi Liu, Zhoukai Zhao, Xiangjun Fan, Benyu Zhang, Yixin Chen
arXiv Machine Learning
Sep 22

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

The paper introduces CS-MoE, a Transformer architecture that shares experts across layers to reduce inter‑layer parameter redundancy. By combining layer‑independent experts with a globally shared expert pool, CS‑MoE allows elastic control over token‑level parameter activation and computational cost. Experiments show that CS‑MoE achieves lower perplexity than equal‑scale dense Transformers while activating only 55% of parameters, and its performance scales with the number of activated experts, approaching MoE performance within a fixed FLOPs budget.

By Dian Jiao, Jiaxin Duan, Shuai Zhao, Jiabing Leng, Yiran Zhang, Feng Huang