arXiv Machine Learning

The Value of Finite Observation in Positive-Data Learning of Multiple Context-Free Languages

arXiv:2605. 11644v3 Announce Type: replace-cross Abstract: Positive data can show that two tuple occurrences share a successful sentence context without certifying that they are safely interchangeable.

arXiv Machine Learning
Sep 4

Relative Prime Factorization and Finite-State Presentations under Fixed Finite-Monoid Observation

The paper investigates exact factorization and canonical presentations of languages relative to a fixed finite‑monoid observation. It shows that unique factorization does not guarantee a finite relative presentation property (FRP) by presenting a 36‑element quotient with infinite valid prime‑return rules, and introduces the stronger finite‑state relative presentation property (FSRP). The authors further define prime‑target left‑division determinism (PTLD), prove its implications for factorization and rule bounds, and provide efficient learning algorithms for the canonical PTLD presentation and FSRP controller.

By Takayuki Kuriyama
arXiv AI
Jun 16

The Faithfulness Gap: Certifying Semantic Equivalence Between Natural-Language and Formal Mathematical Statements

arXiv:2606. 16541v1 Announce Type: new Abstract: Autoformalization, translating natural-language mathematics into formal proof assistants, is bottlenecked not by translation fluency but by \emph{faithfulness}: a formal statement can typecheck and be provable, yet still encode a different theorem than the source intended.

By Noor Islam S. Mohammad, Tamim Sheikh
arXiv Machine Learning
Jun 25

Space-Efficient Language Generation in the Limit

arXiv:2606. 25777v1 Announce Type: cross Abstract: We initiate a resource-aware theory of \textit{language generation in the limit} under the minimal constraint of space efficiency.

By Nicolas Flammarion, Chirag Pabbaraju, Hristo Papazov, Miltiadis Stouras, Ola Svensson
arXiv AI
Aug 12

On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis

arXiv:2601. 05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at the intersection of Algorithmic Information Theory (AIT) and Machine Learning (ML) of great interest.

By Hector Zenil
arXiv Machine Learning
Aug 27

Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events

The paper proposes treating the ‘unit’—a persistent referent that multiple events may refer to—as an explicit primitive in machine learning tasks. It formalizes supervised learning as learning a pair of a tokenizer that generates a contextual unit token and a shared response law that uses this token, thereby distinguishing homogeneous from heterogeneous worlds. The work also introduces concepts such as unit abduction and trusted resolvers to handle cases where unit identity is unresolved.

By Heyang Gong
arXiv Machine Learning
Aug 28

Algorithmic Principles For Multiclass Learning Are Hard To Come By: Limits of Regularization and Proper Learning

The paper investigates fundamental limits of algorithmic principles in multiclass learning, specifically proper learning and regularization. It shows that learning cannot always be reduced to proper learning even with an enlarged hypothesis class, that proper learners may need a sublinear number of errors that can be arbitrarily large, and that regularization (SRM or local) is not universally sufficient. The authors also provide a positive theory giving sufficient conditions for SRM learnability and a characterization via integrability of revealed preferences.

By Julian Asilis, Shaddin Dughmi, Vatsal Sharan, Alec Sun, Shang-Hua Teng, Chang Wang
arXiv Machine Learning
Jul 15

Language Identification with Succinct Machine-Independent Traces

arXiv:2607. 12443v1 Announce Type: cross Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negative results for its original formulation.

By Moses Charikar, Jon Kleinberg, Chirag Pabbaraju