arXiv Machine Learning By Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

Read the original on arXiv Machine Learning →

arXiv:2607. 16184v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.