Information-Theoretic Foundations for Machine Learning
arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.
arXiv:2606. 04045v1 Announce Type: cross Abstract: Representation learning is often described as preserving the information in an input that is relevant for prediction.
arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.
arXiv:2606. 01092v1 Announce Type: cross Abstract: Supervised learning evaluates predictors through their input-output behavior.
arXiv:2606. 19366v1 Announce Type: cross Abstract: Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a hierarchy of abstractions and lifting selected rules back to the signal domain.
The paper proposes treating the ‘unit’—a persistent referent that multiple events may refer to—as an explicit primitive in machine learning tasks. It formalizes supervised learning as learning a pair of a tokenizer that generates a contextual unit token and a shared response law that uses this token, thereby distinguishing homogeneous from heterogeneous worlds. The work also introduces concepts such as unit abduction and trusted resolvers to handle cases where unit identity is unresolved.
arXiv:2609.39525v1 Announce Type: new Abstract: Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intracta...
arXiv:2608. 13510v1 Announce Type: cross Abstract: Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency.
arXiv:2602. 23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning.
arXiv:2605. 08446v3 Announce Type: replace Abstract: Bayesian neural networks are typically trained against the evidence lower bound (ELBO), whose Jensen gap closes only when the variational posterior is exact.
The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.
The paper discusses the data processing inequality (DPI) in statistics, which states that a stochastically modified experiment cannot have a lower Bayes risk than the original. It shows that this classical DPI does not hold for constrained learning problems common in machine learning, where the model class is limited. The authors propose a generalized DPI that applies to constrained Bayes risks, linking it to a set containment condition on a superprediction set, and provide sufficient conditions for this containment.
arXiv:2608.23960v1 Announce Type: cross Abstract: Missing labels are usually regarded as a source of information loss in classification. We study a semi-supervised setting in which the probability of...
arXiv:2603.25579v2 Announce Type: replace-cross Abstract: A key capability of modern neural networks is their capacity to simultaneously learn underlying rules and memorize specific facts or exceptio...