In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update runs until every rollout gets a score.
arXiv:2607. 16241v1 Announce Type: cross Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench.
By Yunxiang Zhang (Xiangjun), Ping Yu (Xiangjun), Jianyu Wang (Xiangjun), Max (Xiangjun), Fan, Julian Reed, Azalia Mirhoseini, Will Su
arXiv:2609.21058v1 Announce Type: cross
Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a front...
By Gaurav Agarwal, Ashish Garg, Isha Singhal
The paper demonstrates that the number of candidates generated during test-time scaling of large language models does not fully capture the system cost. By comparing different generation schedules (e.g., one batched call versus multiple serial calls) while keeping the total candidate count fixed, the authors show that serial calls consume significantly more GPU energy and latency. The study suggests that reporting candidate count alone is insufficient; evaluations should also include generation schedule and GPU-level metrics.
By Mobina Kashaniyan, Ali Jannesari
arXiv:2605. 09708v2 Announce Type: replace-cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in $n$-body problems, multi-field Boltzmann, neighbor-list molecular dynamics, multi-kernel PDE, FFT).
By V\'ictor Gallego
arXiv:2607. 27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness.
By Huihao Jing, Haozhe Cui, Wenbin Hu, Shaojin Chen, Haochen Shi, Changxuan Fan, Yuxuan Liu, Hanyu Yang, Sirui Zhang, Ziyi Chen, Haoran Li, Yangqiu Song
GPU-CFR compiles a fixed game into static dataflow, eliminating per-iteration kernel launches and reducing framework operations by up to 18.1×. On an A100 GPU it achieves 29.8–80.4× speedups over the fastest prior GPU CFR and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the CUDA Graph Replay enables a single graph launch per iteration.
By Boning Li, Longbo Huang
GPU-CFR is a compiler and runtime that transforms any counterfactual regret minimization (CFR) game into a static dataflow representation, eliminating variable kernel launches by precomputing indices, flat arrays, and depth‑level execution blocks. This approach reduces framework operations by up to 18.1× and allows a single CUDA Graph Replay to execute each iteration, yielding 29.8–80.4× speedups over the fastest prior GPU CFR on an A100 and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the optimized path reproduces reference iterates exactly and pays for its overhead within the first solve.
arXiv:2608. 20210v1 Announce Type: cross Abstract: Small language models are usually built like large ones and then squeezed onto a CPU afterwards.
By Christos Koutsiaris
arXiv:2608. 15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps.
By Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang
arXiv:2606. 29119v1 Announce Type: cross Abstract: We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle outer loop over neural-network parameters or structure is worth building.
By Ramchand Kumaresan
arXiv:2607. 11985v1 Announce Type: cross Abstract: Hybrid quantum-classical machine learning workflows repeatedly evaluate many small parametrized circuits during training and model exploration.
By Anton Firc, Martin Pere\v{s}\'ini, Vojt\v{e}ch Mr\'azek, Kamil Malinka, Vojt\v{e}ch Stan\v{e}k, Zbyn\v{e}k Li\v{c}ka, Nouhaila Innan, Walid El Maouaki, Alberto Marchisio, Muhammad Shafique