arXiv:2606. 10493v1 Announce Type: cross Abstract: Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads.
By Wenxin Wang, Yule Hou, Yu Ji, Peng Qu, Youhui Zhang
arXiv:2609.38697v1 Announce Type: cross
Abstract: We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources...
By Matias Parij, Pawan Paudel, Tate Berenbaum, Muthaiah Venkatachalam
arXiv:2605. 21312v2 Announce Type: replace-cross Abstract: Modern LLM serving is no longer homogeneous or monolithic.
By Yicheng Feng, Xin Tan, Yangtao Deng, Yimin Jiang, Yibo Zhu, Hong Xu
arXiv:2504. 08791v3 Announce Type: replace-cross Abstract: On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability.
By Zonghang Li, Tao Li, Wenjiao Feng, Rongxing Xiao, Jianshu She, Hong Huang, Mohsen Guizani, Hongfang Yu, Qirong Ho, Wei Xiang, Xue Liu
The paper introduces Edge0, a streaming mixture‑of‑experts (MoE) inference engine that enables a 35‑billion‑parameter MoE model to run on consumer hardware by predicting routing decisions one token ahead. Edge0 uses a per‑layer prerouter to prefetch the necessary experts from SSD, and an unmerged recovery LoRA trained on the student path to recover quality lost to 4‑bit quantization and routing replacement. On a single 24‑GB machine, Edge0 serves the 35B MoE at 20 tokens per second while keeping peak active memory below 3 GiB, achieving performance close to its fp16 teacher across five public benchmarks.
By Yu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng
arXiv:2606. 19025v1 Announce Type: cross Abstract: Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators.
By Lorenzo Sani, Zeyu Cao, Meghdad Kurmanji, Alex Iacob, Andrej Jovanovic, Yan Gao, Wanru Zhao, Nicholas D. Lane
arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.
By Om Mohite
arXiv:2609.17475v1 Announce Type: new
Abstract: Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit,...
By Yuhua Chen
arXiv:2608.23841v1 Announce Type: cross
Abstract: Single-token autoregressive decode on CPUs is bound by memory bandwidth, not arithmetic: a modern CPU sustains roughly 1 TFLOP/s of compute but only...
By Tom Poperszky
arXiv:2511. 10480v3 Announce Type: replace-cross Abstract: Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed workload execution.
By Changhai Man, Joongun Park, Hanjiang Wu, Huan Xu, Srinivas Sridharan, Tushar Krishna
arXiv:2607. 24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory.
By Dengke Han
arXiv:2607. 23264v1 Announce Type: cross Abstract: Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation.
By Jianwen Xian, Zhiyuan Xu, Yuchen Li, Ziliang Lai, Kang He, Zhen Huang, Aichen Feng, Jinyan Chen, Yilin Zhang, Qinqin Chen, Chengru Song