arXiv Machine Learning

Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation Regime

arXiv:2510. 06028v3 Announce Type: replace Abstract: This paper provides data-dependent bounds on the expected error of the Gibbs algorithm in the overparameterized interpolation regime, where low training errors are also obtained for impossible data, such as random labels in classification.

arXiv Machine Learning
Sep 16

Training Energy-Based Models with Non-MCMC Samplers and Efficient Temperature Estimation

The paper introduces Langevin simulated bifurcation (LSB), a fast, parallel Boltzmann sampler that matches the accuracy of sequential MCMC methods. It also proposes conditional expectation matching (CEM), an efficient technique for estimating the effective temperature of samples from energy‑based models with conditional independence. Building on these, the authors develop sampler adaptive learning (SAL), which adjusts the model temperature to align with the distribution produced by LSB, enabling efficient training of semi‑restricted Boltzmann machines (SRBMs) and outperforming conventional methods on synthetic spin‑glass datasets.

By Kentaro Kubo, Hayato Goto
Hugging Face Trending Papers
Jun 17

Smoothness-Based Derandomization of PAC-Bayes Bounds

We study PAC-Bayes derandomization for smooth loss functions. Our goal is to obtain generalization bounds that hold with high probability for deterministic predictors by exploiting smoothness properties of both the loss and the predictor class.

arXiv AI
Jul 2

Scaling Up Thermodynamic AI Models

arXiv:2607. 00170v1 Announce Type: cross Abstract: Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited.

By Andrew G. Moore
arXiv Statistics ML
4d ago

The double descent and Runge phenomena in overparametrized polynomial interpolation

The paper investigates overparameterized polynomial interpolation across three polynomial bases—Monomial, Chebyshev, and Legendre—using coefficients minimal in the σ^2-norm (and σ^1-norm for the monomial basis). It focuses on equidistant and Chebyshev data points, though many findings hold regardless of sampling specifics. The study draws parallels between the classical Runge phenomenon and the modern double descent phenomenon in machine learning.

By Jason Wein, Stephan Wojtowytsch