arXiv Machine Learning

CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs

arXiv:2608. 00720v1 Announce Type: cross Abstract: Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric.

arXiv Machine Learning
Sep 17

FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining

FAME is an FPGA-based platform that evaluates approximate multipliers directly in hardware, eliminating slow CPU/GPU LUT emulation and reducing evaluation time for DNN inference. It also introduces a pattern-guided retraining method that uses multiplier-specific patterns to recover accuracy losses. Experiments on ResNet‑18 and MobileNetV2 over ImageNet show up to 3.47× faster multiplier evaluation and a 65.5% accuracy improvement over prior retraining approaches.

By Rappy Saha, Nima Amirafshar, Jude Haris, Nima Taherinejad, Jos\'e Cano
arXiv Machine Learning
Jul 1

FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers

arXiv:2606. 31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers.

By Hubert Dymarkowski, Xingjian Fu, Rappy Saha, Jude Haris, Jos\'e Cano
Hugging Face Trending Papers
Jul 9

FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs

Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications. While Field-Programmable Gate Arrays (FPGAs) offer a promising substrate for low-latency inference, conventional FPGA accelerators remain arithmetic-centric, using LUTs primarily as building blocks for numerical operators and peripheral logic.

arXiv Machine Learning
Sep 4

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para-Pipe is a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture for machine‑learning computational graphs on heterogeneous System‑on‑Chip (SoC) platforms. By selectively fine‑tuning parallelism levels across pipeline stages, it navigates the trade‑off between throughput and latency, reducing inter‑processor communication overhead and improving energy efficiency. Evaluation on Amlogic and Black Sesame SoCs shows multiple Pareto‑optimal configurations, with throughput‑optimized setups achieving up to 11.0% better energy efficiency than purely pipelined strategies and 23.3% better than non‑pipelined parallel execution.

By Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra
Hugging Face Trending Papers
Sep 3

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para‑Pipe is a hierarchical mapping framework that combines intra‑ and inter‑stage operator parallelism within a pipelined architecture to optimize deep‑learning inference on heterogeneous System‑on‑Chip (SoC) platforms. By selectively tuning parallelism levels across pipeline stages, it balances throughput and latency while reducing inter‑processor communication overhead. Evaluations on Amlogic and Black Sesame SoCs show Pareto‑optimal configurations, with throughput‑optimized settings achieving up to 11.0 % higher energy efficiency than purely pipelined approaches and 23.3 % over non‑pipelined parallel execution.

arXiv Machine Learning
Sep 4

LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

LeanStream introduces a speculate‑and‑refine streaming framework that enables efficient on‑device inference of large language models by progressively refining computation, loading, and cache‑retention priorities using partial GPU results. This approach allows fine‑grained overlap between GPU execution and storage I/O, avoiding the trade‑offs of existing systems that serialize execution or incur redundant weight fetches. Implemented on mobile and embedded platforms, LeanStream reduces memory usage by 4.8× to 7.5× compared to prior work while improving token generation throughput by 1.6× to 2.1×.

By Renyuan Liu (Richard), Yuyang Leng (Richard), Kaiyan Liu (Richard), Yuzhou Zhong (Richard), Shaohan Hu (Richard), Chun-Fu (Richard), Chen, Peijun Zhao, Heechul Yun, Shuochao Yao
arXiv AI
Aug 18

FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy

arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.

By Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
arXiv AI
Sep 11

DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity

DiffLUT-Net is an FPGA-native neural‑network architecture that uses six‑input lookup tables (LUTs) trained from scratch. The method jointly learns each LUT’s 64 truth‑table entries and the source connections to its six input ports through a differentiable LUT function relaxation and hardware source selection. After training, the learned truth tables and connections are discretized, unused logic is pruned, and the network is exported as synthesizable Verilog, achieving favorable accuracy‑resource trade‑offs across five benchmarks.

By Jiaqi Ye, Xinrui Gong, Jingcun Wang, Olga Kondrateva, Bing Li, Grace Li Zhang