arXiv:2608. 01320v1 Announce Type: cross Abstract: Language generation in the limit is a theoretical framework for studying how a generator can learn to produce new valid strings from a stream of positive examples.
By Ziyi Cai, Shuangping Li, Yiheng Shen, Kangning Wang, Peng Zhang
arXiv:2607. 23361v1 Announce Type: cross Abstract: Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings.
By Debmalya Panigrahi, Fan Wei, Ian Zhang
arXiv:2601. 08648v2 Announce Type: replace-cross Abstract: Recent results in learning a language in the limit have shown that, although language identification is impossible, language generation is tractable.
By Antonios Anastasopoulos, Giuseppe Ateniese, Evgenios M. Kornaropoulos
arXiv:2603. 11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed.
By Giorgio Racca, Michal Valko, Amartya Sanyal
arXiv:2606. 28354v1 Announce Type: cross Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language.
By Irene Strauss, Alexandra Butoi, Ryan Cotterell
arXiv:2606. 25777v1 Announce Type: cross Abstract: We initiate a resource-aware theory of \textit{language generation in the limit} under the minimal constraint of space efficiency.
By Nicolas Flammarion, Chirag Pabbaraju, Hristo Papazov, Miltiadis Stouras, Ola Svensson
arXiv:2607. 12443v1 Announce Type: cross Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with an eye toward variants of the model that might overcome the negative results for its original formulation.
By Moses Charikar, Jon Kleinberg, Chirag Pabbaraju
The paper studies the problem of language generation in the limit, where a learner must produce valid unseen elements from any exhaustive positive presentation of an unknown infinite language. It establishes that generation is possible precisely when each target language admits a finite positive witness such that all targets activated by any finite sample share an infinite common intersection. The authors introduce a separation-width hierarchy to measure the size of compatible witnesses, showing that every level of the hierarchy occurs and that countable families admit singleton witnesses while more complex families require unbounded finite witnesses. The results are formalized and verified in Lean, with the development available on GitHub.
By Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao
arXiv:2606. 14688v1 Announce Type: cross Abstract: AI systems coupled to proof assistants now generate formal mathematics at scale, and the gap between what a checker can verify and what a mathematician would value has become the binding constraint.
By Xiaoyu Li, Andi Han, Dai Shi, Zheng Gao, Jiaojiao Jiang, Junbin Gao
The paper introduces NFA-LM, a polynomial‑time method for generating language model outputs that satisfy nondeterministic finite automaton (NFA) constraints. It leverages recent results that the #NFA counting problem admits a fully polynomial randomized approximation scheme, providing theoretical guarantees under mild assumptions. Experiments demonstrate that NFA‑LM produces high‑quality outputs efficiently while keeping approximation error bounded.
By Jialiang Sun, Kuldeep Meel
arXiv:2602.00612v3 Announce Type: replace
Abstract: Diffusion Large Language Models (dLLMs) have demonstrated promising generative capabilities and are increasingly used to produce formal languages d...
By Yitong Zhang, Yongmin Li, Yuetong Liu, Jia Li, Xiaoran Jia, Zherui Li, Ge Li
The paper introduces PAC‑Private Autoregressive Generation, a method that calibrates noise based on ensemble disagreement across overlapping ‘worlds’ of a private corpus, thereby extending PAC privacy from classification to text generation. By training adapters on a frozen public model and using posterior‑weighted disagreement to add noise only when predictions vary, the approach achieves strong privacy guarantees while preserving most of the fine‑tuning benefit. Experiments on WikiText‑103 with GPT‑2‑small show 74 % of the fine‑tuning gain retained with a per‑token budget of 2⁻³², and membership‑inference success bounded to 51.08 % after one million tokens, outperforming PMixED under matched conditions.
By Mina Mirzadehsarcheshmeh, Amir Keyvan Khandani