arXiv AI

Uncertainty-Aware End-to-End Co-Design of Neural Network Processors: From Training and Mapping to Fabrication

arXiv:2606. 04850v1 Announce Type: cross Abstract: Designing a neural network processor is an end-to-end co-design problem: network architecture and training budget determine the inference workload; hardware mapping decisions determine chip area, latency, and energy; and these characteristics govern fabrication yield and manufacturing cost.

arXiv AI
Jun 26

A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation

arXiv:2606. 20869v2 Announce Type: replace-cross Abstract: We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), which jointly optimizes neural network architectures and their hardware implementations to address the inefficiencies of traditional top-down AI system design flows.

By Selin Yildirim, Yingbing Huang, Deming Chen
arXiv AI
Sep 4

LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks

LevelSyn is a physical-aware logic synthesis framework that uses a level-asynchronous Graph Neural Network to predict high-fidelity gate coordinates by learning the structural and directional semantics of And-Inverter Graphs. It incorporates a level-aligned subgraph partitioning strategy to manage industrial-scale designs and integrates these spatial insights into a new synthesis engine within the Berkeley ABC framework. Experiments on the EPFL benchmark suite show LevelSyn outperforms state-of-the-art methods, achieving an average power reduction of 6.89%, a timing delay improvement of 27.48%, and a 99.59% reduction in design rule check violations.

By Jingyi Zhou, Zhengyuan Shi, Ziyang Zheng, Qiang Xu
arXiv Machine Learning
Jun 8

Amortized Neural Optimization for Pre-Layout Signal Integrity Design Space Exploration using Differentiable Surrogates

arXiv:2606. 07463v1 Announce Type: cross Abstract: Pre-layout design space exploration (DSE) for high-speed signal integrity (SI) analysis is often limited by the computational cost of simulations and iterative optimization algorithms within modern electronic design automation (EDA) workflows.

By Julian With\"oft, Werner John, Emre Ecik, Ralf Br\"uning, J\"urgen G\"otze
Hugging Face Trending Papers
Sep 3

LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks

LevelSyn introduces a physical-aware logic synthesis framework that uses a level-asynchronous Graph Neural Network to predict accurate gate coordinates by learning the structural and directional semantics of And-Inverter Graphs. It incorporates a level-aligned subgraph partitioning strategy to handle large designs and integrates these spatial insights into a new synthesis engine within the Berkeley ABC framework. Experiments on the EPFL benchmark suite show significant improvements, with an average power reduction of 6.89%, a timing delay improvement of 27.48%, and a 99.59% reduction in design rule check violations.

arXiv Machine Learning
Aug 20

Multi-Objective Optimization Under Uncertainty of Part Quality in Fused Filament Fabrication

This study introduces a data‑driven method for multi‑objective optimization in fused filament fabrication (FFF), targeting both geometric accuracy and filament bond quality. Experiments supply part‑quality data, which feed Bayesian neural network models that predict the two objectives while accounting for epistemic and aleatory uncertainties. Using these stochastic predictions, robustness‑based optimization explores nozzle temperature, speed, and layer thickness, producing Pareto surfaces that reveal trade‑offs and are validated through actual part manufacturing.

By Berkcan Kapusuzoglu, Paromita Nath, Matthew Sato, Sankaran Mahadevan, Paul Witherell
Hugging Face Trending Papers
Sep 3

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para‑Pipe is a hierarchical mapping framework that combines intra‑ and inter‑stage operator parallelism within a pipelined architecture to optimize deep‑learning inference on heterogeneous System‑on‑Chip (SoC) platforms. By selectively tuning parallelism levels across pipeline stages, it balances throughput and latency while reducing inter‑processor communication overhead. Evaluations on Amlogic and Black Sesame SoCs show Pareto‑optimal configurations, with throughput‑optimized settings achieving up to 11.0 % higher energy efficiency than purely pipelined approaches and 23.3 % over non‑pipelined parallel execution.