The paper reports a reproducible GPU power benchmark for 18 open‑source LLMs (0.5B–7B parameters) run on a single consumer RTX 4060ti GPU using the Ollama inference engine. Energy metrics such as mean/peak power, total energy per prompt, energy per output token, and throughput were measured, revealing that model architecture and quantization strategy, rather than parameter count alone, drive energy efficiency. The most efficient models were qwen2.5:0.5b and tinyllama:1.1b, while the 7B‑Mistral model consumed up to 8.6× more energy per token, and qwen3.5:0.8b(on) showed unusually high per‑prompt energy due to extended internal reasoning.
By Philipp M. Z\"ahl, Elja Dalipaj, Anika Hennig, Timon Bayer
arXiv:2607. 26571v1 Announce Type: new Abstract: The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems.
By Tina Vartziotis, Rodopi Kosteli, Elli Vartziotis, George Dasoulas, Michael Keckeisen, Konstantinos Skianis, Sotirios Kotsopoulos, Francesca Dominici
TokenPowerSandbox is an evidence‑gated workflow that uses a CPU‑resident projector, brief GPU probes, full‑workload verification, and tamper‑evident provenance to predict energy usage of large language model serving. In experiments on an NVIDIA H100 80GB running Qwen2.5‑7B‑Instruct with vLLM, the method achieved energy MAPE of 6.23% and 7.35% on blind holdout and no‑refit confirmations, with high Spearman rank correlations. A predeclared TTFT gate demonstrated that energy accuracy alone cannot guarantee latency, as it passed at concurrency four but abstained below that level.
By Chenxu Niu
Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient GPUs, but operators currently lack the tools to do so without exhaustively profiling each combination.
arXiv:2607. 02391v1 Announce Type: cross Abstract: Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption.
By Mauricio Fadel Argerich, Jonathan F\"urst, Marta Pati\~no-Mart\'inez
The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure-specific monitoring, limiting its applicability in comparative studies, early-stage system design, and sustainability reporting.