A Fortran General-Purpose Transpiler: Proof of Concept
arXiv:2608. 00130v2 Announce Type: replace-cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains.
JLIR is a Julia-native intermediate representation inspired by MLIR that enables multi-level, dialect-oriented compilation within the Julia ecosystem. It allows Julia programs to be represented before low-level lowering, supports extensible operations and transformation passes via Julia’s language mechanisms, and keeps partially typed programs transformable until concrete types are known. The framework includes built‑in dialects for arithmetic, control flow, functions, structured loops, and memory operations, and can be extended with new domain operations without altering the core system. JLIR was demonstrated by automatically generating JACC kernels for accelerators.
arXiv:2608. 00130v2 Announce Type: replace-cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains.
arXiv:2606. 09930v1 Announce Type: cross Abstract: The boundary between program execution and gradient-based optimization has long limited the use of code itself as a learnable scientific model.
The paper explores using large language models (LLMs) to replace traditional compiler backends, a process termed AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, achieving 0.83x–3.34x the performance of autotuned Triton on a variety of GPUs and ML kernels. The study also extends a PTX verifier to support modern GPU features, highlighting the potential for AI compilers to reduce engineering effort for new hardware.
arXiv:2606. 02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators.
arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.
arXiv:2602. 06142v3 Announce Type: replace-cross Abstract: The phase ordering problem has been a long-standing challenge since the late 1970s, yet it remains an open problem due to having a vast optimization space and an unbounded nature, making it an open-ended problem without a finite solution, one can limit the scope by reducing the number and the length of optimizations.
arXiv:2606. 09213v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are increasingly trained in a wide range of frameworks (SnnTorch, Lava, Norse, and others) each with its own model format.
arXiv:2608. 06791v1 Announce Type: cross Abstract: Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized effort.
arXiv:2606. 04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.
CompileRover is a new optimization framework for virtual machine compilers that uses a tri‑role LLM‑driven collaboration mechanism involving a referee, an advisor, and an operator. It tackles common issues such as redundant computations, inefficient loops, and suboptimal function implementations by applying control‑flow analysis, code‑structure transformations, and dynamic execution pattern recognition. Benchmark evaluations show that CompileRover consistently outperforms existing virtual machine compilers, reducing execution overhead, improving data‑flow consistency, and enhancing overall compiler performance.
arXiv:2607. 20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads.
Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized effort. Even with high-level synthesis (HLS), designers still need extensive hardware expertise to build high-performance accelerators.