arXiv Machine Learning By Hongyu Zhang, Cheng Yan, Xiang Xia, Wuyang Zhang

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

Read the original on arXiv Machine Learning →

arXiv:2608. 04454v1 Announce Type: cross Abstract: Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.