The paper introduces a new complexity measure, based on a restricted class of "layered zipline programs," which is weaker than the previously defined Arithmetic Repetition Complexity (ARC). This measure allows for a prediction algorithm that operates in quasilinear time and polylogarithmic space on highly-structured sequences, offering a more efficient solution at the cost of reduced expressivity compared to ARC. The work highlights a tradeoff between algorithmic efficiency and the breadth of sequences that can be effectively predicted.
By Vanessa Kosoy
The paper investigates integer‑sequence benchmarks from the OEIS by applying a two‑part minimum description length (MDL) learner that searches for P‑recursive recurrences. It finds that MDL difficulty correlates with a combinatorial parameter count, that most sequences fit a recurrence on a prefix but not at full length (the “wilderness” regime), and that language models do not hallucinate in the wilderness but instead hedge, showing that memorisation dominates perceived competence. The study provides a cheap, contamination‑free difficulty signal for OEIS‑derived benchmarks.
By Sabilashan Ganeshan
The article presents a mathematical framework for large language models (LLMs), detailing how text sequences are encoded into tokens, how next‑token prediction architectures are defined, and how these models are trained and deployed for tasks such as summarization, recommendation, software writing, and quantitative problem solving. It emphasizes that the framework relies on basic concepts from information theory, probability, and optimization, yet captures the complex algorithmic structure responsible for LLMs’ empirical successes. The authors argue that this formalism enables the study of accuracy, efficiency, and robustness, and points toward new methodological developments.
By Ricardo Baptista, Andrew Stuart, Son Tran
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
By Jacob Dunefsky, Wes Gurnee, Emmanuel Ameisen
Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language models can do that requires tracking how many tokens remain before a target.
arXiv:2607. 23361v1 Announce Type: cross Abstract: Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings.
By Debmalya Panigrahi, Fan Wei, Ian Zhang
arXiv:2511. 15709v2 Announce Type: replace-cross Abstract: Recent works have shown that tokenisation is NP-complete.
By Violeta Kastreva, Philip Whittington, Dennis Komm, Tiago Pimentel
arXiv:2607. 12443v1 Announce Type: cross Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negative results for its original formulation.
By Moses Charikar, Jon Kleinberg, Chirag Pabbaraju
arXiv:2603.02760v2 Announce Type: replace-cross
Abstract: Diffusion large language models (dLLMs) have recently attracted significant attention for their ability to enhance diversity, controllability...
By Linhao Zhong, Linyu Wu, Wen Wang, Yuling Xi, Chenchen Jing, Jiaheng Zhang, Hao Chen, Chunhua Shen
arXiv:2606. 18856v1 Announce Type: cross Abstract: Sequence labelling, a core task of Natural Language Processing (NLP), consists in assigning each token of an input sentence a label.
By Nicolas Floquet, Joseph Le Roux, Nadi Tomeh
arXiv:2602.13194v3 Announce Type: replace-cross
Abstract: Humans and large language models can predict next letter or word from its prior context much better than random guessing, indicating strong r...
By Weishun Zhong, Doron Sivan, Tankut Can, Mikhail Katkov, Misha Tsodyks
The paper introduces MoirfEolas, a dataset of over 35,000 Irish words annotated with their morphological components, and CríochScore, a metric that measures how well tokenization aligns with these morphological boundaries. Using CríochScore, the authors evaluate common tokenization algorithms and find that the Unigram Language Model best aligns with Irish morphology. They also discuss trade‑offs between morphological alignment, compression, and vocabulary efficiency, offering practical guidance for Irish NLP development.
By Jane Adkins, Abigail Walsh, Brian Davis, Elaine U\'i Dhonnchadha