arXiv Machine Learning

Information-Theoretic Foundations for Machine Learning

arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.

arXiv Machine Learning
Jul 28

Forgetting is Everywhere

arXiv:2511. 04666v4 Announce Type: replace Abstract: A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge as they adapt to new data.

By Ben Sanati, Thomas L. Lee, Trevor McInroe, Aidan Scannell, Esmeralda S. Whitammer, David Abel, Amos Storkey
arXiv AI
Sep 1

Wide Learning: Learning to Reach Evidence

The paper introduces the concept of Wide Learning, which examines how a learner’s internal state can expand its ability to generate informative evidence under fixed resources and primitive affordances. By formalizing effective epistemic reach—defined by learner state, deployment budget, reliability threshold, and evaluation distribution—the authors demonstrate, through a controlled construction, that learning can significantly alter the probability of successfully realizing a diagnostic that was previously unlikely. The study shows that even with identical observable laws, a calibrated learner can achieve perfect diagnostic realization, highlighting the impact of learning on the scope of attainable evidence.

By Junzhou Chen
arXiv Machine Learning
Jun 24

The Degeneracy Distillery

arXiv:2606. 23838v1 Announce Type: new Abstract: When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish.

By T. Lucas Makinen, Deaglan J. Bartlett, Niall Jeffrey, Benjamin D. Wandelt
arXiv Machine Learning
Sep 14

Quantifying the Value of Privileged Information Using a PAC-Bayesian Approach

The paper introduces a PAC‑Bayesian, algorithm‑agnostic framework to quantify the value of privileged information (PI) in Learning Using Privileged Information (LUPI). By comparing the tightest achievable risk bounds with and without PI, the authors derive a training‑time metric that estimates the maximum potential gain from PI without requiring test data. Experiments in supervised and unsupervised settings show a strong correspondence between this metric and actual test‑time performance improvements.

By Vasily Bokov (aQa, Leiden University, The Netherlands, LIACS, Leiden University, Leiden, The Netherlands, Honda Research Institute Europe GmbH, Offenbach, Germany), Sebastian Schmitt (Honda Research Institute Europe GmbH, Offenbach, Germany), Vedran Dunjko (aQa, Leiden University, The Netherlands, LIACS, Leiden University, Leiden, The Netherlands), Hao Wang (aQa, Leiden University, The Netherlands, LIACS, Leiden University, Leiden, The Netherlands)