arXiv AI By Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

Read the original on arXiv AI →

arXiv:2608. 13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 16

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

FlexEE is an early‑exiting framework designed for large language model inference that is constrained by computation and memory, particularly in offloading‑based deployments. It uses layer‑wise exit supervision, self‑speculative decoding over a Top‑K local vocabulary, and dynamic hidden‑state management to enable reliable intermediate‑layer predictions and memory‑aware execution. Experiments on Llama2‑7B and Llama3‑8B show that FlexEE achieves significant speedups—up to 1.27×/3.16× and 1.25×/2.83× respectively—while maintaining minimal accuracy loss.

By Qihu Xie, Ziwei Li, Yi Kang
arXiv AI
Aug 18

S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices

S2-MoE is a self‑speculative decoding framework designed to make Mixture‑of‑Experts (MoE) inference more efficient on edge devices. It reduces verification overhead by using routing‑aware adaptive speculative expansion, improves verification efficiency with reuse‑aware expert gating, and aligns draft and target execution through shared context. Implemented in llama.cpp, S2‑MoE delivers up to 5.3× speedup (≈2.0× on average) over standard autoregressive decoding across various MoE models and datasets on edge hardware.

By Haochen Huang, Shengxuan Qiu, Meng Li