arXiv Machine Learning

Scalable Derivative Gaussian Processes via Exact Gradient Reduction

arXiv:2606. 02909v1 Announce Type: cross Abstract: Gradient observations can substantially improve Gaussian process (GP) surrogates, particularly in high-dimensional settings where function evaluations are expensive.

arXiv Machine Learning
Jul 27

gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points

arXiv:2512. 06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists.

By Marcus M. Noack, Mark D. Risser, Hengrui Luo, Vardaan Tekriwal, Ronald J. Pandolfi
arXiv Machine Learning
Jul 21

Stochastic Dimension Zeroth-Order Estimator: Stable and Memory-Efficient Training of PINNs

arXiv:2603. 24002v3 Announce Type: replace Abstract: Physics-Informed Neural Networks (PINNs) for high-dimensional and high-order partial differential equations (PDEs) are primarily constrained by the $\mathcal{O}(d^k)$ spatial derivative complexity and the $\mathcal{O}(P)$ memory overhead of backpropagation (BP).

By Zhangyong Liang, Huanhuan Gao
arXiv Machine Learning
Sep 3

GRADSOLVE: fast exact gradients for ODE ensembles on GPUs

GRADSOLVE is an open‑source JAX library that provides fast, exact reverse‑mode gradients for low‑dimensional ordinary differential equation (ODE) ensembles on NVIDIA GPUs. It records the accepted steps of an adaptive solver and differentiates a fixed‑step replay, yielding the exact discrete adjoint at a lower computational cost than traditional checkpointed methods. Benchmarks show that GRADSOLVE’s forward kernel is 2.8× faster than DiffEqGPU.jl, and its gradient computation is 5.6–14.1× faster than Diffrax’s checkpointed adjoint while maintaining matched forward‑state accuracy across multiple GPU generations.

By Alessio Spurio Mancini
arXiv Machine Learning
Sep 4

No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels

The paper introduces Finite-Library Input-Warped Bayesian Optimization (FLIWBO), a method that selects input warps from a finite library to adapt the geometry used by Gaussian‑process Bayesian optimization. FLIWBO maintains high‑probability convergence guarantees while improving sample efficiency on problems where raw coordinates poorly match the objective’s geometry, such as log‑scaled hyperparameters or localized peaks. Experiments on synthetic benchmarks, Fashion‑MNIST hyperparameter tuning, and a 20‑dimensional multi‑agent system design demonstrate that FLIWBO‑UCB outperforms raw‑coordinate GP‑UCB and other methods with regret guarantees, especially under misspecified geometry.

By Edvin Ketabati Augustinsson, Robert A. Bridges
arXiv Machine Learning
1d ago

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.

By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)