Hugging Face Trending Papers

Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

Read the original on Hugging Face Trending Papers →

Mixture-of-Experts (MoE) models enable model scaling while maintaining low inference-time compute by activating only a subset of experts per token. However, conventional routing relies on a fixed top-k selection, forcing the model to spend the same compute regardless of how many experts are relevant.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.