arXiv Machine Learning By Zongfei Li

Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts

Read the original on arXiv Machine Learning →

arXiv:2608. 08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.