arXiv Machine Learning

LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load

arXiv:2603. 23640v2 Announce Type: replace-cross Abstract: Deploying large language models on-device for always-on personal agents demands sustained inference from hardware tightly constrained in power, thermal envelope, and memory.

arXiv AI
6d ago

Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware

The paper reports a reproducible GPU power benchmark for 18 open‑source LLMs (0.5B–7B parameters) run on a single consumer RTX 4060ti GPU using the Ollama inference engine. Energy metrics such as mean/peak power, total energy per prompt, energy per output token, and throughput were measured, revealing that model architecture and quantization strategy, rather than parameter count alone, drive energy efficiency. The most efficient models were qwen2.5:0.5b and tinyllama:1.1b, while the 7B‑Mistral model consumed up to 8.6× more energy per token, and qwen3.5:0.8b(on) showed unusually high per‑prompt energy due to extended internal reasoning.

By Philipp M. Z\"ahl, Elja Dalipaj, Anika Hennig, Timon Bayer
arXiv Machine Learning
Jun 24

EnerInfer: Energy-Aware On-Device LLM Inference

arXiv:2606. 23001v1 Announce Type: cross Abstract: On-device LLM inference is increasingly attractive for privacy-preserving, reliable, and cost-effective deployment, yet its energy and thermal costs remain a critical bottleneck.

By Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin, Debayan Roy, Yutao Liu, Yu Peng, Ning Jia, Haibo Chen
arXiv AI
Jun 16

The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution

arXiv:2605. 27599v2 Announce Type: replace-cross Abstract: Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being targeted for edge deployment, with NVIDIA, Dell, HP, ASUS, MSI, Acer, and Gigabyte all shipping GB10-based desktop AI systems in 2026.

By Deepak Panigrahy, Aakash Tyagi
arXiv Machine Learning
Sep 14

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.

By \'Edouard Gu\'egain, Tristan Coignion
arXiv Machine Learning
5d ago

HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference

HybridInfer is a thermal‑aware reinforcement‑learning router that selects among on‑device, edge, and cloud large language model tiers based on a phone’s thermal headroom and query complexity. Trained offline with a Q‑learning policy, it balances quality, latency, cost, and a thermal penalty, and includes a locality bonus that encourages on‑device execution. In real Android tests on 210 prompts, the learned router outperformed two hand‑tuned heuristics in quality while maintaining the lowest cost, and it proved more reliable and faster than always‑on‑device inference, especially for long queries.

By Simran Koul
arXiv AI
Jun 30

KernelSight-LM: A Kernel-Level LLM Inference Simulator

arXiv:2606. 28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets.

By Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt