arXiv Machine Learning By Tianyang Zhu

From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference

Read the original on arXiv Machine Learning →

arXiv:2607. 28097v1 Announce Type: new Abstract: Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.

arXiv Machine Learning
Sep 14

Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference

Dynamic Expert Quantization (DynaExq) is a runtime-aware mixed-precision serving system designed for single‑GPU Mixture‑of‑Experts (MoE) inference under a hard high‑bandwidth memory (HBM) envelope. It treats the problem as an online, budget‑constrained precision allocation task, keeping the most frequently used experts at higher precision while relegating the rest to low‑precision fallbacks. By estimating expert hotness from router traces and asynchronously promoting or demoting experts, DynaExq maintains a fully materialized expert set during the forward pass, improving accuracy and throughput compared to static post‑training quantization and offloading/prefetch baselines. whyItMatters":"DynaExq enables efficient deployment of large MoE models on memory‑limited GPUs by dynamically allocating precision based on runtime expert usage, thereby reducing memory footprint and latency while boosting accuracy and throughput."

By Kexin Chu, Dawei Xiang, Zixu Shen, Yiwei Yang, Zecheng Liu, Wei Zhang
arXiv Machine Learning
Sep 25

Paging the Experts: A Reproducible Characterization of Flash-Backed MoE Inference on iPhone

The paper introduces Routide, a Swift/MLX runtime that runs a quantized Qwen3.6-35B-A3B model on iPhone by keeping expert weights on device storage and a byte‑budgeted subset in memory. It evaluates cache‑policy effects, showing that a 512 MiB LRU cache yields 0.00% demand hits while a 576 MiB LRU reaches 38.58% hits across five 128‑token workloads, indicating that capacity limits depend on policy and workload. The study also reports memory footprints, thermal events, and power estimates, demonstrating that flash‑backed MoE inference is feasible within bounded resources but has measurable limitations.

By Musa Shams