Hugging Face Trending Papers

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

Read the original on Hugging Face Trending Papers →

Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-based methods aggregate routing statistics across all tokens to determine merge compatibility.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.