arXiv:2606. 06656v1 Announce Type: new Abstract: We study parallel Continuous Local Search (CLS) as a solution approach for Boolean satisfiability problems with symmetric pseudo-Boolean (PB) constraints.
By Cody J Christopher, Charles Gretton
arXiv:2606. 08638v1 Announce Type: cross Abstract: Recent research has developed practical, parallelizable first-order methods for large scale linear programming, but performance is highly dependent on hyperparameter selection.
By Siddharth Prasad, Dravyansh Sharma
The paper introduces a fine‑grain GPU implementation of the partition phase of the Generalized Partition Crossover (GPX) for large‑scale Traveling Salesman Problem (TSP) instances. By reformulating GPX partitioning as a graph‑parallel problem with coalesced memory layouts, ghost‑node transformations, and connected‑component analysis, the authors parallelize key operations such as union of parent tours, splitting of degree‑four vertices, deletion of common edges, and component identification using CUDA. Experiments on instances from 10,000 to 2 million cities show speedups between 48× and 625× over a naive sequential CPU implementation while significantly reducing memory overhead.
By Swetha Varadarajan, Darrell Whitley
The paper introduces a GPU-resident, batched Levenberg–Marquardt solver that efficiently optimizes constants in tree-based genetic programming for symbolic regression. By using reverse-mode automatic differentiation to assemble per-tree Jacobians in a single backward sweep, the solver’s per-iteration cost becomes independent of the number of constants per tree, achieving up to 510,000 trees per second on an NVIDIA A100. Integrated into EvoGP, the solver enables end-to-end search that recovers governing equations on 10 of 18 constructed problems, a significant improvement over stock EvoGP.
By Hao Mao, Xu Tony Liu, Shuai Lu, Peng Zhao, Wenzheng Jiang, Yuntian Chen
The paper explores using large language models (LLMs) to replace traditional compiler backends, a process termed AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, achieving 0.83x–3.34x the performance of autotuned Triton on a variety of GPUs and ML kernels. The study also extends a PTX verifier to support modern GPU features, highlighting the potential for AI compilers to reduce engineering effort for new hardware.
By Fran\c{c}ois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.
By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim