Towards Data Science

How Much of a Data Science Workflow Can Run on a GPU Today? Part 1: Accelerating Data Preparation

Exploring GPU acceleration with cuDF, cudf. pandas, and the Polars GPU Engine The post How Much of a Data Science Workflow Can Run on a GPU Today?

arXiv AI
Jun 2

How Much Progress Has There Been in NVIDIA Datacenter GPUs?

arXiv:2601. 20115v3 Announce Type: replace-cross Abstract: As the role of modern Graphics Processing Units (GPUs) becomes increasingly essential for several computing tasks, analyzing their past and current progress is paramount for determining future constraints on scientific research.

By Emanuele Del Sozzo, Martin Fleming, Kenneth Flamm, Neil Thompson
arXiv AI
Sep 12

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

GPU-CFR compiles a fixed game into static dataflow, eliminating per-iteration kernel launches and reducing framework operations by up to 18.1×. On an A100 GPU it achieves 29.8–80.4× speedups over the fastest prior GPU CFR and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the CUDA Graph Replay enables a single graph launch per iteration.

By Boning Li, Longbo Huang
Hugging Face Trending Papers
Sep 10

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

GPU-CFR is a compiler and runtime that transforms any counterfactual regret minimization (CFR) game into a static dataflow representation, eliminating variable kernel launches by precomputing indices, flat arrays, and depth‑level execution blocks. This approach reduces framework operations by up to 18.1× and allows a single CUDA Graph Replay to execute each iteration, yielding 29.8–80.4× speedups over the fastest prior GPU CFR on an A100 and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the optimized path reproduces reference iterates exactly and pays for its overhead within the first solve.