arXiv Machine Learning

Hierarchical Solomonoff Induction: An Unbounded Machine Learning Model

arXiv:2608. 01005v1 Announce Type: new Abstract: Solomonoff Induction, or SolInd, provides an ideal unbounded model of a priori sequence prediction but cannot naturally describe extrapolation from a given training dataset, as performed by Large Language Models.

arXiv AI
Aug 12

On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis

arXiv:2601. 05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at the intersection of Algorithmic Information Theory (AIT) and Machine Learning (ML) of great interest.

By Hector Zenil
arXiv AI
Sep 21

Large Language Models As Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near Without Symbolic Model Synthesis

The paper argues that Large Language Models (LLMs) do not function as Solomonoff induction estimators because their training objectives—cross‑entropy, negative log‑likelihood, and next‑token prediction—optimize fit to a supplied conditional distribution rather than a program‑weighted universal mixture. It further contends that additional computation alone does not transform these models into optimal predictors without external hyper‑parameter or architectural changes. The authors suggest that neurosymbolic machine learning, exemplified by models such as Fable and Astra, represents a shift toward symbolic model synthesis, moving beyond purely statistical LLMs.

By Hector Zenil, Abicumaran Uthamacumaran, Luan Ozelim
arXiv Machine Learning
Sep 1

Adversarial Online Classification with a Preview

arXiv:2608. 29503v1 Announce Type: new Abstract: Worst-case online classification is governed by sequential complexity, such as Littlestone dimension, and can be impossible even for statistically simple classes, such as thresholds of VC dimension one.

By Roi Livni, Sahil Singla
arXiv Machine Learning
Sep 24

Rolling Conformal Prediction in Sequential Model Training

Rolling Conformal Prediction (rolling‑CP) is a distribution‑free predictive inference method designed for sequential model training. It calibrates each incoming observation against the current predictor and incorporates it into future training, eliminating the need for data splitting. For exchangeable data, rolling‑CP guarantees marginal coverage with a universal factor‑two bound, and for i.i.d. streams it provides high‑probability training‑conditional validity over time, improving to the target level under stability conditions.

By Chen Cheng, Ruiting Liang, Rina Foygel Barber
arXiv Machine Learning
Jun 25

Space-Efficient Language Generation in the Limit

arXiv:2606. 25777v1 Announce Type: cross Abstract: We initiate a resource-aware theory of \textit{language generation in the limit} under the minimal constraint of space efficiency.

By Nicolas Flammarion, Chirag Pabbaraju, Hristo Papazov, Miltiadis Stouras, Ola Svensson
arXiv Statistics ML
Sep 7

Reconciling Universal and Uniform Learning with $Q$-Aggregation

The paper investigates regression with bounded responses, comparing two learning frameworks: model selection aggregation, which requires improper algorithms to achieve minimax excess risk, and universal learning, where empirical risk minimization suffices for exponential learning rates. For finite hypothesis classes, the authors show that the $Q$-aggregation estimator simultaneously attains minimax optimal tails and exponential universal rates, while other common estimators fail to do so. For countably infinite classes, they prove an inherent trade‑off between exponential universal and minimax uniform rates, resolved by combining optimal algorithms from each framework via $Q$-aggregation.

By Mikael M{\o}ller H{\o}gsgaard, Patrick Rebeschini, Tobias Wegel
Hugging Face Trending Papers
Aug 13

Bagging Robustly Learns VC Classes with Linear Sample Complexity

We revisit the problem of learning predictors robust to adversarial examples at test-time. We prove that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro (2019).