The paper argues that large language models cannot achieve perfect reliability for any task, even with unlimited scale. It establishes that each generative task has an inherent reliability ceiling set by how much output uncertainty can be resolved from observable context, with a resolvable part that can be improved by more context and a subjective part tied to task ambiguity. The authors derive a scaling law showing that performance is limited by the scarcer resource—either training data or model capacity—and explain how this law explains phenomena such as retrieval augmentation and catastrophic forgetting.
By Subhabrata Majumdar
The paper introduces a new measure of generative‑process diversity for language models, using Normalised Compression Distance on raw outputs after controlling for permutation effects. Across 38 models, this metric uncovers population structure that semantic similarity misses and predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability. The authors argue that higher generative‑process diversity reduces correlated failures in multi‑model systems, offering a practical tool for safety‑relevant applications.
By Ross Tieman, Evan Markou
arXiv:2606. 15327v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have demonstrated strong scaling capacity as alternatives to autoregressive language models.
By Keyue Jiang, Yuxiang Wang, Yanan Zhao, Xiang Yu, Qifang Zhao, Bohan Tang, Baojian Zhou, Yanghua Xiao, Lin Qu, Xiaoxiao Xu
The paper argues that traditional semantic similarity fails to capture the true diversity of language models. It introduces a new metric—generative‑process diversity—measured via Normalised Compression Distance on raw outputs, which reveals hidden population structure among 38 models. This metric predicts lower correlated failures across ten benchmark families, independent of semantic similarity or model capability.
arXiv:2602. 17743v2 Announce Type: replace Abstract: In-context learning (ICL) allows large language models to adapt to new tasks from a few examples without updating their parameters.
By Di Zhang, Ningxu Zhang, Zimeng Liu
arXiv:2409. 02228v2 Announce Type: replace Abstract: When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change?
By Eric Zhang, Leshem Choshen, Jacob Andreas
arXiv:2606. 24752v1 Announce Type: new Abstract: The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning.
By J. Fernando Hernandez-Garcia, Tom\'as Figliolia, Beren Millidge
arXiv:2606. 19857v1 Announce Type: cross Abstract: Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model.
By Jiayi Zhu, Haoxuan Peng, Junxi Wang, Liang Ke, Chen Zhang, Linfeng Zhang
The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning. Although this phenomenon has been known for decades, it has mostly been studied in older, relatively small architectures and rarely in natural-language domains.
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
By Kushal Chakrabarti
arXiv:2608. 15448v1 Announce Type: cross Abstract: Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever.
By Nicolas Zucchet, Hyun Dong Lee, Scott Linderman
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.