arXiv AI By Xuqi Zhu, Huaizhi Zhang, JunKyu Lee, Jiacheng Zhu, Chandrajit Pal, Sangeet Saha, Klaus D. McDonald-Maier, Xiaojun Zhai

Mitigating scalability challenges in LUT-based neural networks via pruning optimisations

Read the original on arXiv AI →

arXiv:2407. 02362v3 Announce Type: replace-cross Abstract: Modern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computational cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining

FAME is an FPGA-based platform that evaluates approximate multipliers directly in hardware, eliminating slow CPU/GPU LUT emulation and reducing evaluation time for DNN inference. It also introduces a pattern-guided retraining method that uses multiplier-specific patterns to recover accuracy losses. Experiments on ResNet‑18 and MobileNetV2 over ImageNet show up to 3.47× faster multiplier evaluation and a 65.5% accuracy improvement over prior retraining approaches.

By Rappy Saha, Nima Amirafshar, Jude Haris, Nima Taherinejad, Jos\'e Cano
arXiv AI
Sep 11

DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity

DiffLUT-Net is an FPGA-native neural‑network architecture that uses six‑input lookup tables (LUTs) trained from scratch. The method jointly learns each LUT’s 64 truth‑table entries and the source connections to its six input ports through a differentiable LUT function relaxation and hardware source selection. After training, the learned truth tables and connections are discretized, unused logic is pruned, and the network is exported as synthesizable Verilog, achieving favorable accuracy‑resource trade‑offs across five benchmarks.

By Jiaqi Ye, Xinrui Gong, Jingcun Wang, Olga Kondrateva, Bing Li, Grace Li Zhang