A Study of Parallel Continuous Local Search
arXiv:2606. 06656v1 Announce Type: new Abstract: We study parallel Continuous Local Search (CLS) as a solution approach for Boolean satisfiability problems with symmetric pseudo-Boolean (PB) constraints.
arXiv:2606. 06641v1 Announce Type: new Abstract: We present Accelerated Fourier SAT (AFSAT), a GPU-accelerated solver for pseudo-Boolean satisfiability based on continuous local search (CLS).
arXiv:2606. 06656v1 Announce Type: new Abstract: We study parallel Continuous Local Search (CLS) as a solution approach for Boolean satisfiability problems with symmetric pseudo-Boolean (PB) constraints.
arXiv:2606. 08638v1 Announce Type: cross Abstract: Recent research has developed practical, parallelizable first-order methods for large scale linear programming, but performance is highly dependent on hyperparameter selection.
The paper introduces a fine‑grain GPU implementation of the partition phase of the Generalized Partition Crossover (GPX) for large‑scale Traveling Salesman Problem (TSP) instances. By reformulating GPX partitioning as a graph‑parallel problem with coalesced memory layouts, ghost‑node transformations, and connected‑component analysis, the authors parallelize key operations such as union of parent tours, splitting of degree‑four vertices, deletion of common edges, and component identification using CUDA. Experiments on instances from 10,000 to 2 million cities show speedups between 48× and 625× over a naive sequential CPU implementation while significantly reducing memory overhead.
The paper introduces a GPU-resident, batched Levenberg–Marquardt solver that efficiently optimizes constants in tree-based genetic programming for symbolic regression. By using reverse-mode automatic differentiation to assemble per-tree Jacobians in a single backward sweep, the solver’s per-iteration cost becomes independent of the number of constants per tree, achieving up to 510,000 trees per second on an NVIDIA A100. Integrated into EvoGP, the solver enables end-to-end search that recovers governing equations on 10 of 18 constructed problems, a significant improvement over stock EvoGP.
The paper explores using large language models (LLMs) to replace traditional compiler backends, a process termed AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, achieving 0.83x–3.34x the performance of autotuned Triton on a variety of GPUs and ML kernels. The study also extends a PTX verifier to support modern GPU features, highlighting the potential for AI compilers to reduce engineering effort for new hardware.
The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.
arXiv:2607. 20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms.
arXiv:2605. 26092v4 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) and Vision Transformers (ViTs) on edge devices is significantly constrained by memory limitations and the critical timing bottlenecks introduced by dense Multiply-Accumulate (MAC) arrays.
arXiv:2606. 07574v1 Announce Type: cross Abstract: Manifold-constrained hyper-connections (mHCs) have recently been proposed as a principled extension of hyper-connections, where the residual mixing matrices are constrained to be doubly stochastic via projection onto the Birkhoff polytope.
arXiv:2608. 15143v1 Announce Type: new Abstract: Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems.
arXiv:2606. 08976v1 Announce Type: new Abstract: LLM-based RTL generation and reasoning is a promising direction for hardware design automation.
arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.