arXiv:2608. 05303v1 Announce Type: cross Abstract: On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications.
By Sangwoo Ha, Hyunwoo Seo, Yurim Jo, Youngjin Moon, Hoi-Jun Yoo
arXiv:2608.05926v2 Announce Type: replace-cross
Abstract: Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM infer...
By Guanqiao Qu, Shuo Chen, Qian Chen, Kin K. Leung, Xianhao Chen
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs.
arXiv:2608. 05926v1 Announce Type: cross Abstract: Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks.
By Guanqiao Qu, Shuo Chen, Qian Chen, Kin K. Leung, Xianhao Chen
arXiv:2607. 24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory.
By Dengke Han
Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel. Existing draft-verify methods use binary decisions: accept or fully recompute.
arXiv:2608. 13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands.
By Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal
FlexEE is an early‑exiting framework designed for large language model inference that is constrained by computation and memory, particularly in offloading‑based deployments. It uses layer‑wise exit supervision, self‑speculative decoding over a Top‑K local vocabulary, and dynamic hidden‑state management to enable reliable intermediate‑layer predictions and memory‑aware execution. Experiments on Llama2‑7B and Llama3‑8B show that FlexEE achieves significant speedups—up to 1.27×/3.16× and 1.25×/2.83× respectively—while maintaining minimal accuracy loss.
By Qihu Xie, Ziwei Li, Yi Kang
arXiv:2606. 12243v1 Announce Type: cross Abstract: Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel.
By Yuchen Xian, Yang He, Yunqiu Xu, Yi Yang
SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.
By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T
arXiv:2608. 10362v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps.
By Eunjeong Kim, Yeong Jun Jeon, Myeonggyun Han