Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs.
arXiv:2608.05926v2 Announce Type: replace-cross
Abstract: Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM infer...
By Guanqiao Qu, Shuo Chen, Qian Chen, Kin K. Leung, Xianhao Chen
S2-MoE is a self‑speculative decoding framework designed to make Mixture‑of‑Experts (MoE) inference more efficient on edge devices. It reduces verification overhead by using routing‑aware adaptive speculative expansion, improves verification efficiency with reuse‑aware expert gating, and aligns draft and target execution through shared context. Implemented in llama.cpp, S2‑MoE delivers up to 5.3× speedup (≈2.0× on average) over standard autoregressive decoding across various MoE models and datasets on edge hardware.
By Haochen Huang, Shengxuan Qiu, Meng Li
arXiv:2609.17193v1 Announce Type: new
Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed...
By Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li
The paper introduces COMLLM, a generative framework that combines Group Relative Policy Optimization with a Look‑Ahead Collaborative Simulation to enable multi‑turn reasoning for task offloading in Mobile Edge Computing. By performing multi‑step Monte Carlo rollouts that jointly model server queue dynamics, COMLLM incorporates long‑term system evolution into its reward design, achieving near‑optimal latency and improved load‑balancing fairness. The framework demonstrates zero‑shot scalability to larger network topologies, outperforming supervised fine‑tuning, deep reinforcement learning, and heuristic baselines without requiring retraining.
By Ning Yang, Chuangxin Cheng, Haijun Zhang
arXiv:2608. 02031v1 Announce Type: cross Abstract: This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints.
By Ngoc Hung Nguyen, Bjorn Landfeldt