Hugging Face Trending Papers

Harness Engineering for LLM-Driven GPU Kernel Generation

Read the original on Hugging Face Trending Papers →

Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be reliably constrained, validated, profiled, and selected. This paper presents a harness-centered system for LLM-driven GPU kernel optimization in the MLSys 2026 FlashInfer AI Kernel Generation Contest on NVIDIA Blackwell B200 GPUs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
4d ago

KLineage: Recovering the Missing When of Kernel Optimization by Deoptimizing Experts

arXiv:2605.28213v2 Announce Type: replace Abstract: LLM-based agents are increasingly used to generate GPU kernels, but they often struggle to determine when an optimization is sound because its requ...

By Shuoming Zhang, Qiuchu Yu, Ruiyuan Xu, Chenjing Zhang, Junjie Peng, Zhicheng Xie, Yangyu Zhang, Xiyu Shi, Ying Liu, Guangli Li, Xiaobing Feng, Huimin Cui, Xingjun Zhang, Jiacheng Zhao
arXiv AI
Sep 16

Ave: Guiding Agentic GPU Optimization Using Data-Flow Invariants

arXiv:2604.18616v2 Announce Type: replace-cross Abstract: LLM coding agents can generate correct GPU kernels, but their performance still trails expert libraries. Reaching peak throughput requires co...

By Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan