Information-Theoretic Limits of Reliability and Scaling in Language Models
arXiv:2607. 14112v1 Announce Type: cross Abstract: Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale.
arXiv:2608. 15798v1 Announce Type: new Abstract: Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to.
arXiv:2607. 14112v1 Announce Type: cross Abstract: Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale.
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
arXiv:2607. 18454v1 Announce Type: cross Abstract: Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling.
arXiv:2607. 20286v1 Announce Type: cross Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt.
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem.
arXiv:2601. 05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at the intersection of Algorithmic Information Theory (AIT) and Machine Learning (ML) of great interest.
arXiv:2607. 18292v1 Announce Type: cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 08961v1 Announce Type: cross Abstract: Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language.
arXiv:2608. 01320v1 Announce Type: cross Abstract: Language generation in the limit is a theoretical framework for studying how a generator can learn to produce new valid strings from a stream of positive examples.
arXiv:2606. 30372v1 Announce Type: new Abstract: Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias.
arXiv:2604. 22167v2 Announce Type: replace-cross Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale.