arXiv AI

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

arXiv:2511. 07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure.

arXiv Machine Learning
Sep 14

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.

By \'Edouard Gu\'egain, Tristan Coignion
arXiv Computer Vision
Aug 25

SemanticXR: Low Power and Real-time Queryable Semantic Mapping with an Object-Level Device-Cloud Architecture

SemanticXR is a device‑cloud system that enables real‑time, open‑vocabulary semantic mapping and querying for XR applications while respecting power, bandwidth, and memory limits. By treating semantically identifiable objects as first‑class units, the system coordinates communication, execution, and memory across device and server, achieving a 2.2× faster server‑side mapping latency and keeping upstream bandwidth below 2.5 Mbps. On the device, an object‑level sparse local map with incremental updates delivers sub‑100 ms query latency for up to 10,000 objects, supports tens of thousands of objects within a 500 MB footprint, and adds only about 2 % to idle power.

By Rahul Singh, Devdeep Ray, Connor Smith, Sarita Adve
arXiv AI
Aug 20

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

The paper presents a method for distributing large language model inference across multiple Intel AI PCs by splitting the model into pipeline shards, each pre‑compiled into an OpenVINO graph. Three key techniques—beam_idx Gather to enable GPU optimizations, speculative decoding on stateful models, and interleaved micro‑batching—allow a two‑node Llama 3.1 8B INT4 pipeline to serve two users at 1.79× the throughput of a single‑node model, while a four‑node deployment can run a 70B model that no single PC can hold. The authors provide code, benchmark logs, and reproduction scripts on GitHub.

By Tate Berenbaum, Muthaiah Venkatachalam
arXiv AI
Sep 15

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

The paper presents a carbon‑aware routing framework for function‑calling in large language models that distributes queries across a three‑tier edge‑cloud architecture. A lightweight k‑NN predictor estimates accuracy, delay, and power for each edge tier, and real‑time grid carbon intensity is used to route queries to the lowest‑emission tier that can execute them. Experiments on state‑of‑the‑art benchmarks show the framework matches cloud‑level accuracy while cutting operational carbon emissions by an average of four times.

By Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos
arXiv AI
Aug 18

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.

By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv Machine Learning
Sep 11

Optimizing AI Inference Across the Deployment Stack

The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.

By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel