arXiv:2603. 20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult.
By Yingwei Zheng, Cong Li, Shaohua Li, Yuqun Zhang, Zhendong Su
arXiv:2606. 20373v1 Announce Type: cross Abstract: Large Language Models (LLMs) show promise for code compilation tasks, but applying them to runtime performance tuning is difficult due to complex microarchitectural effects and noisy runtime measurements.
By Zepeng Li, Jie Ren, Zhanyong Tang, Jie Zheng, Zheng Wang
arXiv:2606. 02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators.
By Taras Sereda, Burak Bartan, Ankita Nayak, Tom St. John, Natalie Serrino, Zain Asgar
CUDA‑Harness is a framework that enables the generation and optimization of CUDA kernels directly from natural language. It introduces Intermediate‑Structured Generation to bridge high‑level semantics with low‑level kernel code, uses Synthesis‑Based Verification to mitigate reward hacking by providing isolated test data, and employs Feedback‑Adaptive Evolution to prioritize correctness while improving performance. Experiments show the approach generalizes across different large language models, hardware platforms, and even supports C‑to‑CUDA transpilation.
By Qi Fan, An Zou, Yehan Ma
arXiv:2608. 03983v1 Announce Type: cross Abstract: Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation.
By Hailong Jiang, Feng Yu, Emran Hossain, Jianfeng Zhu, Mengfei Ren, Qiang Guan, Chunwei Xia
arXiv:2604.18616v2 Announce Type: replace-cross
Abstract: LLM coding agents can generate correct GPU kernels, but their performance still trails expert libraries. Reaching peak throughput requires co...
By Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan