arXiv:2608. 05499v1 Announce Type: cross Abstract: Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices.
By Sadegh Jafari, Mohiuddin Bilwal, Fan Zhou, Brian Gelder, Ali Jannesari
arXiv:2606. 07819v1 Announce Type: new Abstract: Recently, the efficiency of Large Language Models (LLMs) deployment has become a critical concern in practical applications.
By Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi, Phuong Hoai Ha
arXiv:2608. 15693v1 Announce Type: new Abstract: Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation.
By Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
The paper introduces RAMP, a method for robust adaptive mixed‑precision quantization of vision models on edge CPUs. It evaluates 13 sensitivity metrics across four neural networks, finding that Jensen‑Shannon Divergence consistently identifies layers that can be safely quantized. Using K‑Means clustering on these metrics, RAMP achieves near‑lossless accuracy with an average 1.81× speed‑up, while cautioning against excluding low‑speed‑up layers that can fragment the computational graph.
By David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. Howeve...
arXiv:2506.11784v2 Announce Type: replace
Abstract: Vision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-wid...
By Guang Liang, Xinyao Liu, Jianxin Wu