arXiv AI By Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao

Shape Mutating Expert Compression:LorExperts and BTExperts

Read the original on arXiv AI →

arXiv:2608. 07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

ITC-MoE introduces an Importance-guided Token-aware Compression framework for Mixture-of-Experts Diffusion Language Models. It combines Adaptive Tucker Compression, which uses activation and gradient importance to jointly factorize expert weights and allocate ranks, with Token-aware Compensation and Routing that applies low‑rank adjustments to hot tokens and limits expert candidates for cold tokens. The method achieves significant reductions in computation and storage while maintaining generation quality, exemplified by a 30% compression budget that preserves 96.33% accuracy on MultiArith and delivers up to a 7.22× speedup.

By Lianjun Liu, Shipeng Li, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong