arXiv AI

On the Instance Hardness as a Decision Criterion in TinyML Systems

The paper discusses TinyML, which deploys machine learning on devices with limited memory and computing resources. It introduces a preliminary study using the tree depth prune instance hardness method to control thresholds in TinyML systems. The results suggest that adjusting these thresholds can reduce energy consumption while maintaining classification quality.

arXiv AI
Aug 25

Power-Performance Characterization of TinyML Systems

arXiv:2608.21646v1 Announce Type: cross Abstract: TinyML systems are enabling machine learning (ML) inference at the edge. However, there is little quantitative analysis of such systems. This paper p...

By Yujie Zhang, Dhananjaya Wijerathne, Zhaoying Li, Tulika Mitra
arXiv Machine Learning
Sep 24

Reliable Federated TinyML Deployment for IoT Security

The paper explores how to combine Federated Learning with TinyML model compression techniques—such as knowledge distillation, structured pruning, and quantization—to create lightweight, privacy‑preserving intrusion detection systems for IoT devices. It evaluates these strategies within a federated training pipeline and finds that training stability is crucial; a server‑coordinated cosine learning‑rate schedule boosts Attack Recall from 46.7% to 93.85% while still allowing significant model compression and efficient edge deployment.

By Younsoo Park, Seokhyoen Bae, Shasi Kumar Ramachandran Prabhu, Suman Saha, Peilong Li
arXiv Machine Learning
Sep 14

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.

By \'Edouard Gu\'egain, Tristan Coignion
arXiv Machine Learning
Jun 30

Harvesting AI Computation at the Edge via Generic Approximation

arXiv:2606. 29518v1 Announce Type: cross Abstract: With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge.

By Yihan Wang, Huiru Yan, Luxin Zhang, Long Cheng, Weiwei Chen, Ying Wang, Lei Zhang, Cheng Liu, Huawei Li