arXiv Machine Learning

The Normalized Maximum Likelihood for Regular Non-Smooth Models: Measure-Theoretic Foundations and Geometric Sampling

The paper develops a rigorous framework for computing the Normalized Maximum Likelihood (NML) codelength for regular path‑differentiable Lipschitz (PDL) estimators, which include non‑smooth models such as Lasso and Sparse SVMs. By leveraging geometric measure theory and a novel Propose‑and‑Project Metropolis‑Hastings sampler, the authors provide a method to exactly evaluate the stochastic complexity for these non‑smooth estimators and demonstrate its scalability to high‑dimensional settings. The study shows that the exact NML criterion can match cross‑validation performance while being more data‑efficient, offering a theoretically grounded alternative for model selection in modern machine learning.

arXiv Machine Learning
Jun 19

Fisher-Geometric Sharpness and the Implicit Bias of SGD toward Flat Minima

arXiv:2606. 20469v1 Announce Type: new Abstract: A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative.

By Md Sakir Ahmed, Kumaresh Sarmah, Hemen Dutta
arXiv Machine Learning
Sep 7

The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures

The paper investigates how many linear samples are needed to learn Lipschitz operators under Gaussian measures. It establishes both lower and upper bounds on the Hermite polynomial approximation error and shows that the minimal worst‑case error cannot converge algebraically with the number of samples. However, if the covariance operator of the Gaussian measure decays rapidly, convergence rates arbitrarily close to any algebraic rate can be achieved.

By Ben Adcock, Michael Griebel, Gregor Maier
arXiv Machine Learning
Sep 4

A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations

arXiv:2609. 03129v1 Announce Type: cross Abstract: Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks.

By Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
arXiv Statistics ML
Aug 26

A Non-asymptotic Analysis for Learning and Applying a Preconditioner in MCMC

The paper presents a non‑asymptotic analysis of Markov chain Monte Carlo (MCMC) algorithms that learn and apply a preconditioner based on either the target covariance or the expected Hessian of the target potential. It compares the finite‑time computational costs of these preconditioned schemes with unpreconditioned counterparts, providing guarantees for algorithms such as the Unadjusted Langevin Algorithm (ULA) and the proximal sampler. The analysis relies on a contraction assumption in the Wasserstein‑2 distance to formalize approximate independence and bridge modern MCMC theory with classical effective sample size heuristics.

By Max Hird, Florian Maire, Jeffrey Negrea