Paging the Experts: A Reproducible Characterization of Flash-Backed MoE Inference on iPhone
Read the original on arXiv Machine Learning →The paper introduces Routide, a Swift/MLX runtime that runs a quantized Qwen3.6-35B-A3B model on iPhone by keeping expert weights on device storage and a byte‑budgeted subset in memory. It evaluates cache‑policy effects, showing that a 512 MiB LRU cache yields 0.00% demand hits while a 576 MiB LRU reaches 38.58% hits across five 128‑token workloads, indicating that capacity limits depend on policy and workload. The study also reports memory footprints, thermal events, and power estimates, demonstrating that flash‑backed MoE inference is feasible within bounded resources but has measurable limitations.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.