arXiv Machine Learning By Maoliang Li, Haojing Chen, Jiayu Chen, Zihao Zheng, Xinhao Sun, Hailong Zou, Xiang Chen

MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers

Read the original on arXiv Machine Learning →

arXiv:2606. 15615v1 Announce Type: new Abstract: Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation across timesteps.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.