Hugging Face Trending Papers

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

Read the original on Hugging Face Trending Papers →

GPU-CFR is a compiler and runtime that transforms any counterfactual regret minimization (CFR) game into a static dataflow representation, eliminating variable kernel launches by precomputing indices, flat arrays, and depth‑level execution blocks. This approach reduces framework operations by up to 18.1× and allows a single CUDA Graph Replay to execute each iteration, yielding 29.8–80.4× speedups over the fastest prior GPU CFR on an A100 and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the optimized path reproduces reference iterates exactly and pays for its overhead within the first solve.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 12

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

GPU-CFR compiles a fixed game into static dataflow, eliminating per-iteration kernel launches and reducing framework operations by up to 18.1×. On an A100 GPU it achieves 29.8–80.4× speedups over the fastest prior GPU CFR and 14–258× over the LiteEFG CPU implementation for large games. The compiled representation alone delivers 2.2–51.1× acceleration on eight CPU threads, while the CUDA Graph Replay enables a single graph launch per iteration.

By Boning Li, Longbo Huang
arXiv Machine Learning
Sep 23

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.

By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim