arXiv Computation and Language

JLIR: A Julia-Native MLIR-Inspired Intermediate Representation with Automatic JACC Kernel Extraction

JLIR is a Julia-native intermediate representation inspired by MLIR that enables multi-level, dialect-oriented compilation within the Julia ecosystem. It allows Julia programs to be represented before low-level lowering, supports extensible operations and transformation passes via Julia’s language mechanisms, and keeps partially typed programs transformable until concrete types are known. The framework includes built‑in dialects for arithmetic, control flow, functions, structured loops, and memory operations, and can be extended with new domain operations without altering the core system. JLIR was demonstrated by automatically generating JACC kernels for accelerators.

arXiv AI
4d ago

AI as a Compiler: Compiling Triton kernels without the Triton compiler

The paper explores using large language models (LLMs) to replace traditional compiler backends, a process termed AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, achieving 0.83x–3.34x the performance of autotuned Triton on a variety of GPUs and ML kernels. The study also extends a PTX verifier to support modern GPU features, highlighting the potential for AI compilers to reduce engineering effort for new hardware.

By Fran\c{c}ois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv AI
Jun 15

Protean Compiler: An Agile Framework to Drive Fine-grain Phase Ordering

arXiv:2602. 06142v3 Announce Type: replace-cross Abstract: The phase ordering problem has been a long-standing challenge since the late 1970s, yet it remains an open problem due to having a vast optimization space and an unbounded nature, making it an open-ended problem without a finite solution, one can limit the scope by reducing the number and the length of optimizations.

By Amir H. Ashouri, Shayan Shirahmad Gale Bagi, Kavin Satheeskumar, Tejas Srikanth, Jonathan Zhao, Ibrahim Saidoun, Ziwen Wang, Bryan Chan, Tomasz S. Czajkowski
arXiv AI
Aug 10

HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation

arXiv:2608. 06791v1 Announce Type: cross Abstract: Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized effort.

By Yuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura, Le Chen, Venkatram Vishwanath, Zhenman Fang, Caiwen Ding
arXiv Machine Learning
Sep 17

CompileRover: Revolutionizing Virtual Machine Compiler Optimization with a Tri-Role LLM-Driven Framework

CompileRover is a new optimization framework for virtual machine compilers that uses a tri‑role LLM‑driven collaboration mechanism involving a referee, an advisor, and an operator. It tackles common issues such as redundant computations, inefficient loops, and suboptimal function implementations by applying control‑flow analysis, code‑structure transformations, and dynamic execution pattern recognition. Benchmark evaluations show that CompileRover consistently outperforms existing virtual machine compilers, reducing execution overhead, improving data‑flow consistency, and enhancing overall compiler performance.

By Mingqiao Mo, Yunlong Tan, Hao Zhang