arXiv AI

WARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI

WARD is a runtime‑adaptive Vision Transformer framework designed for edge AI that combines channel‑wise subnetwork partitioning, reliability‑aware continual learning, and dynamic operating‑mode scheduling. It operates two physically isolated subnetworks across four modes—Full‑Precision, Low‑Power, High‑Reliability, and Adaptive—to balance computational cost and fault tolerance while maintaining uninterrupted inference. Implemented on a lightweight FPGA accelerator with minimal area overhead, WARD achieves a network‑level failure rate of 1.79% under high Bit Error Rates and supports rapid mode transitions within a few clock cycles.

arXiv AI
Jul 14

Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles

arXiv:2607. 10942v1 Announce Type: cross Abstract: Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints.

By Ashiyana Abdul Majeed, Mahmoud Meribout, Neethu Joseph, Abel Kidane Haile, Mohammad Abdullah Al Faruque
arXiv Computer Vision
Sep 14

Adaptive AI: Energy Efficient Multi-exit TinyML on Intelligent Vision Systems at the Edge

The paper presents a novel multi‑exit computational scheme for TinyML on an ultra‑low‑power GAP9 SoC, adding confidence‑based gating points to a MobileNetV2 CNN for ImageNet‑100. By allowing inference to stop early, the approach cuts average MAC operations by 41 % (from 313 MMAC to 185 MMAC), reduces inference time by 29 % (49 ms to 35 ms), and saves 24 % in energy (2.1 mJ to 1.6 mJ per frame) with only a ~1 % drop in accuracy. Compared to a state‑of‑the‑art adaptive CNN on the same hardware, the method more than doubles computational efficiency, raising MAC/cycle from 8.1 to 17.2.

By Luca Crupi, Lorenzo Lamberti, Alessandro Giusti, Daniele Palossi
arXiv Machine Learning
Jul 1

FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers

arXiv:2606. 31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers.

By Hubert Dymarkowski, Xingjian Fu, Rappy Saha, Jude Haris, Jos\'e Cano
arXiv Machine Learning
Sep 24

RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models

The paper introduces RAMP, a method for robust adaptive mixed‑precision quantization of vision models on edge CPUs. It evaluates 13 sensitivity metrics across four neural networks, finding that Jensen‑Shannon Divergence consistently identifies layers that can be safely quantized. Using K‑Means clustering on these metrics, RAMP achieves near‑lossless accuracy with an average 1.81× speed‑up, while cautioning against excluding low‑speed‑up layers that can fragment the computational graph.

By David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones
arXiv AI
Sep 17

BLADE: ReliaBle Dynamic Hardware-Aware SNN-ANN Boundary SeLection for Event-BAseD Object DEtection

The paper introduces BLADE, a reliability‑aware method for selecting the boundary between spiking and artificial neural network components in hybrid event‑based object detectors. It jointly optimizes boundary placement and early‑exit settings for reliability, accuracy, execution time, and energy, using fault injection to guide design. Experiments show a 0.691 mAP@0.5 with 15.82 mJ energy when the ANN exits early, and that protecting a single floating‑point exponent bit eliminates catastrophic failures while a fully SNN configuration retains 96.5% reliability under severe faults.

By Mahdi Taheri, Alwin Paul