arXiv Machine Learning By Yu Zhang

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes

Read the original on arXiv Machine Learning →

arXiv:2608. 07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.