arXiv Machine Learning

Energy-efficient operation of neural operators for virtual sensing

The paper studies how sharing spatial computation can lower the energy consumption of neural operator updates in virtual sensing tasks. By comparing compiler freezing, explicit trunk reuse, and graph replay, the authors find energy reductions ranging from about 1% at low request rates to 20% at high rates, with a 22–22.5% saving over eager execution in a 15 W mode. The work also highlights how different neural operator architectures (DeepONet and FNO) and overheads such as preparation and worker replacement affect the overall energy profile.

arXiv Machine Learning
Jun 2

Real-Time Sensing of Inaccessible Physical Fields via an Edge-Deployable Hardware-Portable Graph Neural Operator

arXiv:2604. 01802v2 Announce Type: replace Abstract: Real-time inference of inaccessible interior physical fields from sparse boundary observations is a fundamental but unresolved problem in scientific machine learning, with direct relevance to safety-critical monitoring across many engineering applications.

By William Howes, Jason Yoo, Kazuma Kobayashi, Subhankar Sarkar, Farid Ahmed, Souvik Chakraborty, Syed Bahauddin Alam
arXiv Machine Learning
Aug 26

Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Edge-Deployable Virtual Sensing

The paper introduces the Sparse-Activation-ReLU (SAR) layer, a single‑step neural operator that promotes activation sparsity without surrogate‑gradient training and is compatible with event‑based computing. In a trunk‑based NOMAD architecture, SAR improves the combined Latency‑Error‑Energy (LEE) metric by over fivefold compared to Variable Spiking Neuron (VSN) and Leaky Integrate‑and‑Fire (LIF) models. Additional techniques such as synthetic knowledge distillation, a ReLU‑based spiking loss, and graph‑neighbor thresholding further reduce LEE and L2 error on the Heat Exchanger dataset, advancing energy‑efficient virtual sensing for edge deployment.

By William Howes, Farid Ahmed, Syed Bahauddin Alam
arXiv Machine Learning
Jun 30

Harvesting AI Computation at the Edge via Generic Approximation

arXiv:2606. 29518v1 Announce Type: cross Abstract: With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge.

By Yihan Wang, Huiru Yan, Luxin Zhang, Long Cheng, Weiwei Chen, Ying Wang, Lei Zhang, Cheng Liu, Huawei Li
arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv Machine Learning
Sep 7

Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled Hardware

The paper presents an empirically calibrated, state‑aware Dynamic Voltage and Frequency Scaling (DVFS) scheduler that eliminates thermal throttling on a passively cooled Raspberry Pi 5 during sustained YOLOv8n inference. By using time‑domain guards, absolute temperature bounds, and derivative triggers, the scheduler outperforms a temperature‑only baseline with a 6.8% higher frame rate and 1.9% less energy per frame, and it surpasses an actively cooled reference in energy efficiency. The study also identifies that the passive operating envelope closes at ambient temperatures above 27 °C, where nonlinear leakage undermines DVFS control, and demonstrates that correct scheduling can make mechanical cooling unnecessary within the mapped envelope.

By Aayush Marasini, Zhaoxian Zhou
arXiv AI
Sep 12

DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management

The paper proposes DCO, a dynamic cache orchestration scheme for multi-core AI accelerators that uses application-aware policies and dataflow information to guide cache replacement, bypass decisions, and thrashing mitigation. Using a cycle-accurate simulator, the authors demonstrate up to 1.80× speedup over conventional cache architectures and validate the approach with an analytical model and RTL implementation. The design occupies 0.064 mm² on a 15 nm process and operates at 2 GHz, showing that a shared system-level cache can simplify programming while boosting performance for large language model workloads.

By Zhongchun Zhou, Chengtao Lai, Yuhang Gu, Wei Zhang