arXiv:2608. 04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults.
By Mohammad Hasan Ahmadilivani, Sven-Markus Loorits, Jaan Raik
TreeFI is a value‑aware statistical fault‑injection technique for FP32 single‑bit faults in deep neural network activations and weights. It partitions each layer’s value distribution into intervals with similar expected bit‑flip behavior using regression trees, then allocates injections across these intervals based on their relevance for failure‑rate estimation. This stratified approach preserves target confidence and error margins while dramatically reducing the required injection budget—up to 72.1× for activations and 11.2× for weights compared to existing baselines.
By Noam Bires, Marcello Traiola, Angeliki Kritikakou, Elisa Fromont
arXiv:2405. 01741v4 Announce Type: replace-cross Abstract: Reliability of AI systems is a fundamental concern for the successful deployment and widespread adoption of AI technologies.
By Xun Jiao, Fred Lin, Harish D. Dixit, Joel Coburn, Sajin Nair, Abhinav Pandey, Han Wang, Venkat Ramesh, Jianyu Huang, Daniel Moore, Sriram Sankar
arXiv:2608.29598v1 Announce Type: new
Abstract: We present the first systematic study of the resilience of text-to-video (T2V) diffusion models under random hardware-level faults. While T2V models ar...
By Zachary Coalson, A M Aahad, Stella Doehring, Zane Ma, Sanghyun Hong
arXiv:2603. 22770v2 Announce Type: replace-cross Abstract: The deployment of deep neural networks (DNNs) in safety-critical edge environments necessitates robustness against hardware-induced bit-flip errors.
By Alan T. L. Bacellar, Sathvik Chemudupati, Shashank Nag, Allison Seigler, Priscila M. V. Lima, Felipe M. G. Fran\c{c}a, Lizy K. John
arXiv:2609.28099v1 Announce Type: cross
Abstract: Deep vision systems remain vulnerable to corruption, occlusion, and distribution shift despite strong benchmark performance. Existing reliability met...
By Anoushka Harit, Rehan Zuberi, William Prew, Florian Markowetz
arXiv:2607. 19623v1 Announce Type: cross Abstract: We characterize per-bit-position fault sensitivity in ML inference across 16 workloads -- spanning transformer-based models and attention-free CNNs -- and across three floating-point formats.
By Muhammad Husnain Mubarik, Karthik Mohan Kumar, Pedro Antonio Pena, Keshavan Varadarajan, Kunal Tyagi
The paper introduces REQAP, a reliability‑aware quantized weight packing technique for systolic‑array DNN accelerators. It uses a sensitivity‑driven mixed‑precision quantization to assign layer‑wise bit‑widths, a deterministic register‑level packing strategy for SIMD‑within‑a‑register execution, and selective bit‑level protection that replicates critical MSBs into unused register space. Experiments on AlexNet, VGG‑11, and ResNet‑18 show up to 62% memory reduction, 56% fewer MAC operations, and improved accuracy resilience under fault injection compared to baseline and fully protected models.
By Mahdi Taheri, Samira Nazari, Mubassher Ansari, Ali Azarpeyvand, Mohsen Afsharchi, Maksim Jenihhin, Christian Herglotz
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
arXiv:2606. 09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact.
By Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou
WARD is a runtime‑adaptive Vision Transformer framework designed for edge AI that combines channel‑wise subnetwork partitioning, reliability‑aware continual learning, and dynamic operating‑mode scheduling. It operates two physically isolated subnetworks across four modes—Full‑Precision, Low‑Power, High‑Reliability, and Adaptive—to balance computational cost and fault tolerance while maintaining uninterrupted inference. Implemented on a lightweight FPGA accelerator with minimal area overhead, WARD achieves a network‑level failure rate of 1.79% under high Bit Error Rates and supports rapid mode transitions within a few clock cycles.
By Mahdi Taheri, Pramit Kumar Bhaduri, Mohammad Masoumi, Ali Mahani
arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.
By Dipankar Sarkar