The paper introduces REQAP, a reliability‑aware quantized weight packing technique for systolic‑array DNN accelerators. It uses a sensitivity‑driven mixed‑precision quantization to assign layer‑wise bit‑widths, a deterministic register‑level packing strategy for SIMD‑within‑a‑register execution, and selective bit‑level protection that replicates critical MSBs into unused register space. Experiments on AlexNet, VGG‑11, and ResNet‑18 show up to 62% memory reduction, 56% fewer MAC operations, and improved accuracy resilience under fault injection compared to baseline and fully protected models.
By Mahdi Taheri, Samira Nazari, Mubassher Ansari, Ali Azarpeyvand, Mohsen Afsharchi, Maksim Jenihhin, Christian Herglotz
arXiv:2606. 09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact.
By Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou
arXiv:2603. 22770v2 Announce Type: replace-cross Abstract: The deployment of deep neural networks (DNNs) in safety-critical edge environments necessitates robustness against hardware-induced bit-flip errors.
By Alan T. L. Bacellar, Sathvik Chemudupati, Shashank Nag, Allison Seigler, Priscila M. V. Lima, Felipe M. G. Fran\c{c}a, Lizy K. John
arXiv:2605. 24391v2 Announce Type: replace-cross Abstract: As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference.
By Dahoon Park, Jahyun Koo, Sangwoo Hwang, Jaeha Kung
arXiv:2609.37887v1 Announce Type: new
Abstract: Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds...
By Arian Eamaz, Mojtaba Soltanalian
The paper introduces RAMP, a method for robust adaptive mixed‑precision quantization of vision models on edge CPUs. It evaluates 13 sensitivity metrics across four neural networks, finding that Jensen‑Shannon Divergence consistently identifies layers that can be safely quantized. Using K‑Means clustering on these metrics, RAMP achieves near‑lossless accuracy with an average 1.81× speed‑up, while cautioning against excluding low‑speed‑up layers that can fragment the computational graph.
By David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones