arXiv AI By Peijun Zhu, Ning Yang, Baoliang Tian, Jiayu Wei, Weihao Zhang, Haijun Zhang, Pin Lv

Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression

Read the original on arXiv AI →

arXiv:2510. 02345v4 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.