arXiv AI

GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators

arXiv:2602. 22352v2 Announce Type: replace-cross Abstract: With the continuous growth of neural network scales, low-precision quantization is widely used in edge accelerators.

arXiv AI
2d ago

ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator

ShatterQuant is a hardware-software co-designed framework that enables mixed-precision quantization within individual tensors by assigning different bit-widths to blocks of a weight projection. It couples precision granularity with processing element configuration, allowing each precision to determine an effective block height. The framework includes a hardware-aware post-training method based on block-level standard deviation and weight sensitivity, a ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities, and an evaluation showing 1.5 TOPS, 760 GOPS/$mm^2$ area efficiency, and 2.8 TOPS/W energy efficiency on a TSMC 16nm PDK implementation.

By Mikolaj Walczak, Edward Humes, Chao Fang, Marian Verhelst, Tinoosh Mohsenin
arXiv Machine Learning
Aug 28

Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

The paper presents a PyTorch-based framework for designing and optimizing binarized neural networks, incorporating freezing and pruning mechanisms. It introduces a novel pruning method that uses a global weighting scheme to assess parameter importance across abstraction levels, achieving a 70% pruning rate on VGG11 without sacrificing accuracy—outperforming existing binarized pruning results of 41%. The framework facilitates rapid, reproducible evaluation and prototyping of state‑of‑the‑art binarized network techniques.

By Roan Rubiales, Jean Pierre David
arXiv Machine Learning
Sep 17

FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining

FAME is an FPGA-based platform that evaluates approximate multipliers directly in hardware, eliminating slow CPU/GPU LUT emulation and reducing evaluation time for DNN inference. It also introduces a pattern-guided retraining method that uses multiplier-specific patterns to recover accuracy losses. Experiments on ResNet‑18 and MobileNetV2 over ImageNet show up to 3.47× faster multiplier evaluation and a 65.5% accuracy improvement over prior retraining approaches.

By Rappy Saha, Nima Amirafshar, Jude Haris, Nima Taherinejad, Jos\'e Cano
arXiv AI
Aug 5

A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures

arXiv:2311. 17815v3 Announce Type: replace-cross Abstract: Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performance Computing (HPC) and edge platforms, leading to a wide variety of proposals for specialized deep learning architectures and hardware accelerators.

By Serena Curzel, Fabrizio Ferrandi, Leandro Fiorin, Daniele Ielmini, Cristina Silvano, Francesco Conti, Luca Bompani, Luca Benini, Enrico Calore, Sebastiano Fabio Schifano, Cristian Zambelli, Maurizio Palesi, Giuseppe Ascia, Enrico Russo, Valeria Cardellini, Salvatore Filippone, Francesco Lo Presti, Stefania Perri
arXiv Machine Learning
Jun 30

Physical Analogue Kolmogorov-Arnold Networks based on Reconfigurable Nonlinear-Processing Units

arXiv:2602. 07518v3 Announce Type: replace-cross Abstract: Kolmogorov-Arnold Networks (KANs) shift neural computation from linear layers to learnable nonlinear edge functions, but implementing these nonlinearities efficiently in hardware remains an open challenge.

By Manuel Escudero, Mohamadreza Zolfagharinejad, Sjoerd van den Belt, Nikolaos Alachiotis, Wilfred G. van der Wiel
arXiv Computer Vision
Sep 18

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

The paper introduces MiX, a micro‑inverted‑scaling format that replaces shared exponents with shared mantissas to avoid microscaling collapse in low‑bit vision‑language models. An adaptive dual‑format inference framework (MiX‑MX) maps this format to a custom accelerator, replacing multipliers with shifters. Experiments show 4.5‑bit MiX matches or outperforms NVFP4 accuracy while improving area efficiency by 25 % and delivering 2.3–4.5× speedup with 1.4–2.9× energy savings over the Focus accelerator.

By Yuan Liao, Jae-sun Seo
arXiv AI
Sep 2

HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference

HBQ: Hierarchical Scaling Block Quantization with Hardware‑Efficiency‑Aware Design for Accurate LLM Inference proposes a new block‑quantization scheme that uses large blocks and low‑overhead significand scaling to balance hardware efficiency and accuracy. The authors demonstrate that larger blocks improve efficiency by amortizing dequantization and accumulation costs, while their SIG scaling compensates for the resulting accuracy loss. Experiments on a 28 nm ASIC accelerator show that HBQ achieves up to 4.6× higher area/energy efficiency than state‑of‑the‑art weight‑only quantization, with 1.5–3.0× speedup and 1.6–3.3× system energy reduction over existing BQ methods.

By Chun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal, Mark Anders, Mohamed Abdelfattah, Jae-sun Seo