Mixture of Experts (MoEs) in Transformers
Related stories
A theoretical model for task routing in mixture-of-expert transformers
arXiv:2606. 14398v1 Announce Type: new Abstract: Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed.
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Mixture of Experts Explained
Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
arXiv:2607. 20426v1 Announce Type: cross Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal knowledge or have poor cross-domain generalization.
Training nGPT
arXiv:2608. 01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere.
Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
Mixtral of experts
Stability of Transformers under Layer Normalization
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
Yes, Transformers are Effective for Time Series Forecasting (+ Autoformer)
Symmetry-Aware Transformer Training for Automated Planning
arXiv:2508. 07743v2 Announce Type: replace Abstract: While transformers excel in many settings, their application in the field of automated planning is limited.