FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
Read the original on arXiv AI →FlexEE is an early‑exiting framework designed for large language model inference that is constrained by computation and memory, particularly in offloading‑based deployments. It uses layer‑wise exit supervision, self‑speculative decoding over a Top‑K local vocabulary, and dynamic hidden‑state management to enable reliable intermediate‑layer predictions and memory‑aware execution. Experiments on Llama2‑7B and Llama3‑8B show that FlexEE achieves significant speedups—up to 1.27×/3.16× and 1.25×/2.83× respectively—while maintaining minimal accuracy loss.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.