arXiv Computation and Language By Ning Li, Xinyu Wang, Xin Yuan, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang

TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference

Read the original on arXiv Computation and Language →

The paper introduces TopoCompress, a token compression framework designed for distributed edge Mixture-of-Experts (MoE) inference. It jointly optimizes token compression, expert deployment, GPU-CPU residency, and routing to reduce cross-server communication and resource usage. The method uses a two-timescale alternating optimization, with an online loop compressing low-importance tokens and an offline loop updating expert placement based on accumulated traffic.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jul 21

OrderMoE: An expert similarity driven distributed edge MoE inference

arXiv:2607. 17154v1 Announce Type: cross Abstract: Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over resource-constrained and bandwidth-limited edge infrastructures.

By Xin Yuan, Ning Li, Quan Chen, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang
Hugging Face Trending Papers
Jul 19

OrderMoE: An expert similarity driven distributed edge MoE inference

Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over resource-constrained and bandwidth-limited edge infrastructures. Existing distributed MoE serving methods mainly rely on exact expert placement, caching, replication, or communication scheduling, while overlooking the functional similarity among experts, which provides an opportunity to reduce cross-server token transmission.