Hugging Face Trending Papers

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para‑Pipe is a hierarchical mapping framework that combines intra‑ and inter‑stage operator parallelism within a pipelined architecture to optimize deep‑learning inference on heterogeneous System‑on‑Chip (SoC) platforms. By selectively tuning parallelism levels across pipeline stages, it balances throughput and latency while reducing inter‑processor communication overhead. Evaluations on Amlogic and Black Sesame SoCs show Pareto‑optimal configurations, with throughput‑optimized settings achieving up to 11.0 % higher energy efficiency than purely pipelined approaches and 23.3 % over non‑pipelined parallel execution.

arXiv Machine Learning
Sep 4

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para-Pipe is a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture for machine‑learning computational graphs on heterogeneous System‑on‑Chip (SoC) platforms. By selectively fine‑tuning parallelism levels across pipeline stages, it navigates the trade‑off between throughput and latency, reducing inter‑processor communication overhead and improving energy efficiency. Evaluation on Amlogic and Black Sesame SoCs shows multiple Pareto‑optimal configurations, with throughput‑optimized setups achieving up to 11.0% better energy efficiency than purely pipelined strategies and 23.3% better than non‑pipelined parallel execution.

By Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra
arXiv AI
Aug 5

A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures

arXiv:2311. 17815v3 Announce Type: replace-cross Abstract: Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performance Computing (HPC) and edge platforms, leading to a wide variety of proposals for specialized deep learning architectures and hardware accelerators.

By Serena Curzel, Fabrizio Ferrandi, Leandro Fiorin, Daniele Ielmini, Cristina Silvano, Francesco Conti, Luca Bompani, Luca Benini, Enrico Calore, Sebastiano Fabio Schifano, Cristian Zambelli, Maurizio Palesi, Giuseppe Ascia, Enrico Russo, Valeria Cardellini, Salvatore Filippone, Francesco Lo Presti, Stefania Perri
arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv AI
Jun 26

A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation

arXiv:2606. 20869v2 Announce Type: replace-cross Abstract: We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), which jointly optimizes neural network architectures and their hardware implementations to address the inefficiencies of traditional top-down AI system design flows.

By Selin Yildirim, Yingbing Huang, Deming Chen
Hugging Face Trending Papers
Jun 24

Energy-Efficient CNN Acceleration with MSDF Digit-Serial Arithmetic on FPGA

This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemented on FPGA. While digit-serial arithmetic, particularly most-significant-digit-first (MSDF) techniques, offers a compact hardware footprint, it suffers from initial latency before producing the first output digit.

arXiv AI
Jul 13

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

arXiv:2607. 09385v1 Announce Type: cross Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs).

By Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini