arXiv Machine Learning By Maoliang Li, Haojing Chen, Jiayu Chen, Zihao Zheng, Xinhao Sun, Hailong Zou, Xiang Chen

MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers

Read the original on arXiv Machine Learning →

arXiv:2606. 15615v1 Announce Type: new Abstract: Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation across timesteps.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

ITC-MoE introduces an Importance-guided Token-aware Compression framework for Mixture-of-Experts Diffusion Language Models. It combines Adaptive Tucker Compression, which uses activation and gradient importance to jointly factorize expert weights and allocate ranks, with Token-aware Compensation and Routing that applies low‑rank adjustments to hot tokens and limits expert candidates for cold tokens. The method achieves significant reductions in computation and storage while maintaining generation quality, exemplified by a 30% compression budget that preserves 96.33% accuracy on MultiArith and delivers up to a 7.22× speedup.

By Lianjun Liu, Shipeng Li, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong