arXiv Machine Learning

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

arXiv:2607. 19712v1 Announce Type: new Abstract: In RLHF pipelines, reward scoring blocks policy updates.

arXiv Machine Learning
Sep 18

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

The paper demonstrates that the number of candidates generated during test-time scaling of large language models does not fully capture the system cost. By comparing different generation schedules (e.g., one batched call versus multiple serial calls) while keeping the total candidate count fixed, the authors show that serial calls consume significantly more GPU energy and latency. The study suggests that reporting candidate count alone is insufficient; evaluations should also include generation schedule and GPU-level metrics.

By Mobina Kashaniyan, Ali Jannesari
arXiv AI
Sep 12

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

GPU-CFR compiles a fixed game into static dataflow, eliminating per-iteration kernel launches and reducing framework operations by up to 18.1×. On an A100 GPU it achieves 29.8–80.4× speedups over the fastest prior GPU CFR and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the CUDA Graph Replay enables a single graph launch per iteration.

By Boning Li, Longbo Huang
Hugging Face Trending Papers
Sep 10

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

GPU-CFR is a compiler and runtime that transforms any counterfactual regret minimization (CFR) game into a static dataflow representation, eliminating variable kernel launches by precomputing indices, flat arrays, and depth‑level execution blocks. This approach reduces framework operations by up to 18.1× and allows a single CUDA Graph Replay to execute each iteration, yielding 29.8–80.4× speedups over the fastest prior GPU CFR on an A100 and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the optimized path reproduces reference iterates exactly and pays for its overhead within the first solve.

arXiv Machine Learning
Jul 15

VQCSim: When Does Compile-Once Statevector Simulation Beat Generic Quantum Frameworks?

arXiv:2607. 11985v1 Announce Type: cross Abstract: Hybrid quantum-classical machine learning workflows repeatedly evaluate many small parametrized circuits during training and model exploration.

By Anton Firc, Martin Pere\v{s}\'ini, Vojt\v{e}ch Mr\'azek, Kamil Malinka, Vojt\v{e}ch Stan\v{e}k, Zbyn\v{e}k Li\v{c}ka, Nouhaila Innan, Walid El Maouaki, Alberto Marchisio, Muhammad Shafique