arXiv Machine Learning

Stringological sequence prediction II: Right-to-left automaticity and related complexity measures

arXiv:2607. 17369v1 Announce Type: cross Abstract: In a previous paper, we began the study of sequence prediction algorithms adapted to stringological word complexity measures.

arXiv Machine Learning
Sep 18

Stringological sequence prediction III: layered ziplines and a tradeoff between efficiency and expressivity

The paper introduces a new complexity measure, based on a restricted class of "layered zipline programs," which is weaker than the previously defined Arithmetic Repetition Complexity (ARC). This measure allows for a prediction algorithm that operates in quasilinear time and polylogarithmic space on highly-structured sequences, offering a more efficient solution at the cost of reduced expressivity compared to ARC. The work highlights a tradeoff between algorithmic efficiency and the breadth of sequences that can be effectively predicted.

By Vanessa Kosoy
arXiv Machine Learning
Sep 1

Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence Benchmarks

The paper investigates integer‑sequence benchmarks from the OEIS by applying a two‑part minimum description length (MDL) learner that searches for P‑recursive recurrences. It finds that MDL difficulty correlates with a combinatorial parameter count, that most sequences fit a recurrence on a prefix but not at full length (the “wilderness” regime), and that language models do not hallucinate in the wilderness but instead hedge, showing that memorisation dominates perceived competence. The study provides a cheap, contamination‑free difficulty signal for OEIS‑derived benchmarks.

By Sabilashan Ganeshan
arXiv Machine Learning
Aug 19

Large Language Models: A Mathematical Formulation

The article presents a mathematical framework for large language models (LLMs), detailing how text sequences are encoded into tokens, how next‑token prediction architectures are defined, and how these models are trained and deployed for tasks such as summarization, recommendation, software writing, and quantitative problem solving. It emphasizes that the framework relies on basic concepts from information theory, probability, and optimization, yet captures the complex algorithmic structure responsible for LLMs’ empirical successes. The authors argue that this formalism enables the study of accuracy, efficiency, and robustness, and points toward new methodological developments.

By Ricardo Baptista, Andrew Stuart, Son Tran
arXiv Machine Learning
Jul 28

Hallucination Rates in Language Generation

arXiv:2607. 23361v1 Announce Type: cross Abstract: Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings.

By Debmalya Panigrahi, Fan Wei, Ian Zhang
arXiv Machine Learning
Jul 15

Language Identification with Succinct Machine-Independent Traces

arXiv:2607. 12443v1 Announce Type: cross Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negative results for its original formulation.

By Moses Charikar, Jon Kleinberg, Chirag Pabbaraju
arXiv Computation and Language
Sep 7

MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology

The paper introduces MoirfEolas, a dataset of over 35,000 Irish words annotated with their morphological components, and CríochScore, a metric that measures how well tokenization aligns with these morphological boundaries. Using CríochScore, the authors evaluate common tokenization algorithms and find that the Unigram Language Model best aligns with Irish morphology. They also discuss trade‑offs between morphological alignment, compression, and vocabulary efficiency, offering practical guidance for Irish NLP development.

By Jane Adkins, Abigail Walsh, Brian Davis, Elaine U\'i Dhonnchadha