arXiv AI By Lie Li, Wen Li, Junxiao Shen, Gusheng Hu

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

Read the original on arXiv AI →

arXiv:2608. 15299v1 Announce Type: cross Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.