arXiv AI

PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization

arXiv:2608. 13790v1 Announce Type: cross Abstract: Macro placement significantly affects a chip's post-route performance, power, and area (PPA).

arXiv Machine Learning
Jun 26

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

arXiv:2606. 26453v1 Announce Type: new Abstract: We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating large language model (LLM) code generation with hardware profiler feedback and pluggable bottleneck detection tools.

By Jiading Gai, Shuai Zhang, Kaj Bostrom, Jin Huang, Vihang Patil, Haoyang Fang, Bernie Wang, Huzefa Rangwala, George Karypis
Hugging Face Trending Papers
Jun 8

Order Matters: Unveiling the Hidden Impact of Macro Placement Sequences via Proxy-Guided LLM Evolution

Macro placement is a fundamental step in modern chip physical design, playing a crucial role in determining the solution quality of high-dimensional combinatorial optimization problems. Despite recent advancements in machine learning for spatial coordinate determination, the temporal dimension of placement sequencing remains largely governed by static heuristics.

arXiv Machine Learning
Sep 11

Optimizing AI Inference Across the Deployment Stack

The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.

By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel
arXiv AI
Sep 4

LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks

LevelSyn is a physical-aware logic synthesis framework that uses a level-asynchronous Graph Neural Network to predict high-fidelity gate coordinates by learning the structural and directional semantics of And-Inverter Graphs. It incorporates a level-aligned subgraph partitioning strategy to manage industrial-scale designs and integrates these spatial insights into a new synthesis engine within the Berkeley ABC framework. Experiments on the EPFL benchmark suite show LevelSyn outperforms state-of-the-art methods, achieving an average power reduction of 6.89%, a timing delay improvement of 27.48%, and a 99.59% reduction in design rule check violations.

By Jingyi Zhou, Zhengyuan Shi, Ziyang Zheng, Qiang Xu