Hugging Face Blog

Mixture of Experts (MoEs) in Transformers

arXiv AI
Jul 24

Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations

arXiv:2607. 20426v1 Announce Type: cross Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal knowledge or have poor cross-domain generalization.

By Xinyue Fang, Zhiliang Tian, Zhen Huang, Ziyi Pan, Zhihua Wen, Xi Wang, Quntian Fang, Dongsheng Li
arXiv Machine Learning
Aug 4

Training nGPT

arXiv:2608. 01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere.

By Ilya Loshchilov, Boris Ginsburg
arXiv Machine Learning
Sep 22

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

The paper introduces CS-MoE, a Transformer architecture that shares experts across layers to reduce inter‑layer parameter redundancy. By combining layer‑independent experts with a globally shared expert pool, CS‑MoE allows elastic control over token‑level parameter activation and computational cost. Experiments show that CS‑MoE achieves lower perplexity than equal‑scale dense Transformers while activating only 55% of parameters, and its performance scales with the number of activated experts, approaching MoE performance within a fixed FLOPs budget.

By Dian Jiao, Jiaxin Duan, Shuai Zhao, Jiabing Leng, Yiran Zhang, Feng Huang
arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li