arXiv Machine Learning By Xin Yuan, Ning Li, Quan Chen, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang

OrderMoE: An expert similarity driven distributed edge MoE inference

Read the original on arXiv Machine Learning →

arXiv:2607. 17154v1 Announce Type: cross Abstract: Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over resource-constrained and bandwidth-limited edge infrastructures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 19

OrderMoE: An expert similarity driven distributed edge MoE inference

Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over resource-constrained and bandwidth-limited edge infrastructures. Existing distributed MoE serving methods mainly rely on exact expert placement, caching, replication, or communication scheduling, while overlooking the functional similarity among experts, which provides an opportunity to reduce cross-server token transmission.

arXiv Computation and Language
Sep 23

TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference

The paper introduces TopoCompress, a token compression framework designed for distributed edge Mixture-of-Experts (MoE) inference. It jointly optimizes token compression, expert deployment, GPU-CPU residency, and routing to reduce cross-server communication and resource usage. The method uses a two-timescale alternating optimization, with an online loop compressing low-importance tokens and an offline loop updating expert placement based on accumulated traffic.

By Ning Li, Xinyu Wang, Xin Yuan, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang