arXiv:2604. 01802v2 Announce Type: replace Abstract: Real-time inference of inaccessible interior physical fields from sparse boundary observations is a fundamental but unresolved problem in scientific machine learning, with direct relevance to safety-critical monitoring across many engineering applications.
By William Howes, Jason Yoo, Kazuma Kobayashi, Subhankar Sarkar, Farid Ahmed, Souvik Chakraborty, Syed Bahauddin Alam
The paper introduces the Sparse-Activation-ReLU (SAR) layer, a single‑step neural operator that promotes activation sparsity without surrogate‑gradient training and is compatible with event‑based computing. In a trunk‑based NOMAD architecture, SAR improves the combined Latency‑Error‑Energy (LEE) metric by over fivefold compared to Variable Spiking Neuron (VSN) and Leaky Integrate‑and‑Fire (LIF) models. Additional techniques such as synthetic knowledge distillation, a ReLU‑based spiking loss, and graph‑neighbor thresholding further reduce LEE and L2 error on the Heat Exchanger dataset, advancing energy‑efficient virtual sensing for edge deployment.
By William Howes, Farid Ahmed, Syed Bahauddin Alam
arXiv:2606. 29518v1 Announce Type: cross Abstract: With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge.
By Yihan Wang, Huiru Yan, Luxin Zhang, Long Cheng, Weiwei Chen, Ying Wang, Lei Zhang, Cheng Liu, Huawei Li
arXiv:2606. 20537v1 Announce Type: new Abstract: Mainstream LLM serving systems reuse prefix work mainly through paged or radix key-value (KV) caches.
By Liang Su
arXiv:2607. 29398v1 Announce Type: new Abstract: Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising.
By Zhikang Xie, Xichen Ye, Yifan Wu, Haoshen Yu, Li chenan, Peizhu Gong, Weizhong Zhang, Cheng Jin
arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.
By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
The paper presents an empirically calibrated, state‑aware Dynamic Voltage and Frequency Scaling (DVFS) scheduler that eliminates thermal throttling on a passively cooled Raspberry Pi 5 during sustained YOLOv8n inference. By using time‑domain guards, absolute temperature bounds, and derivative triggers, the scheduler outperforms a temperature‑only baseline with a 6.8% higher frame rate and 1.9% less energy per frame, and it surpasses an actively cooled reference in energy efficiency. The study also identifies that the passive operating envelope closes at ambient temperatures above 27 °C, where nonlinear leakage undermines DVFS control, and demonstrates that correct scheduling can make mechanical cooling unnecessary within the mapped envelope.
By Aayush Marasini, Zhaoxian Zhou
arXiv:2605. 21312v2 Announce Type: replace-cross Abstract: Modern LLM serving is no longer homogeneous or monolithic.
By Yicheng Feng, Xin Tan, Yangtao Deng, Yimin Jiang, Yibo Zhu, Hong Xu
arXiv:2608.30070v1 Announce Type: new
Abstract: Sparse representations are often expected to make models smaller and also reduce inference cost. For Fourier Neural Operators (FNOs), these objectives...
By Abdul Qadir Ibrahim, Martin Burger
The paper proposes DCO, a dynamic cache orchestration scheme for multi-core AI accelerators that uses application-aware policies and dataflow information to guide cache replacement, bypass decisions, and thrashing mitigation. Using a cycle-accurate simulator, the authors demonstrate up to 1.80× speedup over conventional cache architectures and validate the approach with an analytical model and RTL implementation. The design occupies 0.064 mm² on a 15 nm process and operates at 2 GHz, showing that a shared system-level cache can simplify programming while boosting performance for large language model workloads.
By Zhongchun Zhou, Chengtao Lai, Yuhang Gu, Wei Zhang
arXiv:2606. 01839v1 Announce Type: cross Abstract: LLM-based agents resolve a user task through many turns of dependent inference and tool calls, producing a workload whose total cost is unknown when the task arrives.
By Jianru Ding, Ryien Hosseini, Pouya Mahdi Gholami, Mingyuan Xiang, Henry Hoffmann
arXiv:2609.09772v1 Announce Type: new
Abstract: SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spik...
By Ting Liu