arXiv AI

SOLAR: AI-Powered Speed-of-Light Performance Analysis

arXiv:2606. 26383v1 Announce Type: cross Abstract: How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit?

arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv AI
Jul 24

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

arXiv:2607. 20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs.

By Arya Tschand, Charles Hong, Julian Walker, Nina Cai, Shangkun Wang, Suvinay Subramanian, Sundar Dev, Vijay Janapa Reddi, Amir Yazdanbakhsh, Sethu Sankaran
arXiv AI
Jul 24

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

arXiv:2607. 20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms.

By Xue-Jian Gao, Deng Pan, Yueming Su, Jiasheng Li, Bin Du, Fengming Zhu, Chengdi Ma, Junyi Fan, Qichen Liao, Chengqiu Hu, Xinxian Chen, Lingchao Zheng, Jun Li, Jiwei Yang, Yuwei Fan
arXiv AI
Jul 3

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

arXiv:2607. 01590v1 Announce Type: new Abstract: Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints and strict memory hierarchies.

By Junyi Wen, Ruiyan Zhuang, Yongjia Xu, Pengtu Li, Rui Zou, Hongyi Chen, Chingman Wan, Puxu Yang, Wuhui Chen, Yanlin Wang
arXiv AI
Aug 28

PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices

PICasso is an AI‑enabled framework that converts natural‑language specifications into manufacturable silicon photonic integrated circuits (PICs) through a structured pipeline of NL → YAML → GDS, PDK‑aware knowledge injection, automated placement and routing, DRC/LVS validation, and SAX‑based photonic simulation. The authors introduce PIC‑Set, a benchmark of 36 parameterized PIC design tasks, and evaluate several large language models (LLMs) using new metrics such as structural and functional Spec@k, optimization efficiency, and robustness. Across the benchmark, PICasso markedly improves specification satisfaction, achieving up to 92.7% structural Spec@3 and 52% functional Spec@3, while reducing mean insertion loss from 4.98 dB to 3.25 dB through simulation‑guided optimization.

By Deepak Vungarala, Deniz Najafi, Abdulrahman Aljoudi, Zahra Ghanaatian, Navid Khoshavi, Gourav Datta, Arman Roohi, Mahdi Nikdast, Shaahin Angizi
arXiv AI
Aug 19

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

The paper explores energy-aware knowledge distillation for large language models (LLMs) used in software engineering tasks such as clone detection, vulnerability prediction, and code summarization. It shows that the commonly used FLOPs metric does not reliably reflect actual energy consumption, and that using energy-surrogate models during distillation can reduce inference energy by up to 90% and memory usage by 86% with only modest accuracy loss. The study demonstrates that guiding distillation with direct energy estimates improves the sustainability and deployability of LLMs on consumer hardware.

By Enrique Barba Roque, Lu\'is Cruz, Annibale Panichella