arXiv Machine Learning

The Method of Gaps: Exact Expressions for the Generalization Error of Supervised Learning Algorithms

arXiv:2411. 12030v3 Announce Type: replace Abstract: In this paper, the method of gaps, a technique for deriving closed-form expressions in terms of information measures for the generalization error of supervised learning algorithms, is introduced.

arXiv Machine Learning
Sep 25

Machine Unlearning for Gibbs Supervised Learning Algorithms

The paper introduces a method for exact unlearning of Gibbs supervised learning algorithms via a variational formulation based on empirical risk minimization with relative entropy regularization (ERM‑RER). By maximizing the expected empirical risk over the data to be removed while regularizing with relative entropy to the original algorithm, the resulting solution is a new Gibbs probability measure that matches the distribution of an algorithm retrained from scratch on the remaining data. The approach also provides a general framework for reweighting data points in ERM‑RER, allowing for up‑ or down‑weighting to control generalization error or other objectives.

By Yaiza Bermudez, Samir M. Perlaza, I\~naki Esnaola
arXiv Statistics ML
Sep 18

Equivalence Between Nested Gibbs Measures and Log-Linear Combinations of Gibbs Measures

The paper investigates three operations on Gibbs probability measures: renormalization, normalized log-linear combination, and nesting (changing the reference measure). It shows that the measures produced by the second and third operations solve related optimization problems and that, for specific parameters, nesting one Gibbs measure into another is equivalent to log-linearly combining them. This equivalence has practical implications, such as enabling a one-shot federated learning system where clients’ locally trained Gibbs algorithms can be combined on a server to match the performance of a centrally trained Gibbs algorithm.

By Yaiza Bermudez, Samir M. Perlaza, I\~naki Esnaola
arXiv Machine Learning
Sep 11

General Quantification of Covariate and Concept Shifts

arXiv:2609. 11918v1 Announce Type: new Abstract: Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples.

By Hongbo Chen, Li Charlie Xia
Hugging Face Trending Papers
Jul 14

Solution of the Hempel's statistical ambiguity problem and Causal AI

This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are derived from statistical laws. To avoid such predictions, Carl Hempel proposed the Requirement of Maximal Specificity (RMS) for the statistical laws used in the inference.

arXiv Machine Learning
Aug 20

Learning Random Geometric Graphs Drawn in Probabilistic Metric Spaces

The paper introduces a data‑driven method for learning Random Geometric Graphs (RGGs) in probabilistic metric spaces. It defines a distance function based on the cumulative distribution of a disparity variable that captures differences in vertex connectivity and correlation of attached random variables, enabling edges to exist with a specified probability. The approach includes a rejection‑sampling technique for edge probability estimation and a closed‑form posterior for learning the inter‑observable correlation matrix, and it is demonstrated on highly multivariate real datasets.

By Dalia Chakrabarty, Kangrui Wang, Chuqiao Zhang, Ye Liu