arXiv:2607. 19712v1 Announce Type: new Abstract: In RLHF pipelines, reward scoring blocks policy updates.
By Venkata Naga Sai Vishnu Rohit Pulipaka, Anish Katta, Deva Rohit Reddy Peddireddy
arXiv:2607. 16241v1 Announce Type: cross Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench.
By Yunxiang Zhang (Xiangjun), Ping Yu (Xiangjun), Jianyu Wang (Xiangjun), Max (Xiangjun), Fan, Julian Reed, Azalia Mirhoseini, Will Su
arXiv:2609.21058v1 Announce Type: cross
Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a front...
By Gaurav Agarwal, Ashish Garg, Isha Singhal
arXiv:2605. 09708v2 Announce Type: replace-cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in $n$-body problems, multi-field Boltzmann, neighbor-list molecular dynamics, multi-kernel PDE, FFT).
By V\'ictor Gallego
arXiv:2607. 27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness.
By Huihao Jing, Haozhe Cui, Wenbin Hu, Shaojin Chen, Haochen Shi, Changxuan Fan, Yuxuan Liu, Hanyu Yang, Sirui Zhang, Ziyi Chen, Haoran Li, Yangqiu Song
GPU-CFR is a compiler and runtime that transforms any counterfactual regret minimization (CFR) game into a static dataflow representation, eliminating variable kernel launches by precomputing indices, flat arrays, and depth‑level execution blocks. This approach reduces framework operations by up to 18.1× and allows a single CUDA Graph Replay to execute each iteration, yielding 29.8–80.4× speedups over the fastest prior GPU CFR on an A100 and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the optimized path reproduces reference iterates exactly and pays for its overhead within the first solve.