The paper investigates the relationship between a large language model’s internal probability distribution and its verbalized confidence statements. By systematically manipulating training and in‑context data, the authors show that both internal and verbalized probabilities are influenced by distributional and asserted uncertainty in the data. They find that verbalized probabilities align with internal ones beyond what would be expected if they tracked the same sources independently, indicating that verbalized confidence can serve as a probe of the model’s internal distribution.
By Sinead Williamson, Jiaxuan Li, Nick Foti, Russ Webb, Masha Fedzechkina
arXiv:2603.18908v5 Announce Type: replace
Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
By Matt Gorbett, Suman Jana
arXiv:2506. 14003v5 Announce Type: replace Abstract: Machine unlearning (MU) for large language models (LLMs), commonly referred to as LLM unlearning, seeks to remove specific undesirable data or knowledge from a trained model, while maintaining its performance on standard tasks.
By Yiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu, Sijia Liu
arXiv:2606. 07951v1 Announce Type: cross Abstract: Humans increasingly turn to Language Models (LMs) in ways that shape beliefs and drive decisions, including discussing, rewriting, and summarizing information from scientific articles, news, and medical reports.
By Catarina G Belem, Shang Wu, Hongyu Yao, Mark Steyvers, Sameer Singh, Padhraic Smyth
arXiv:2609.34187v2 Announce Type: replace-cross
Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
By Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
arXiv:2609.15533v1 Announce Type: cross
Abstract: Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inheren...
By Tobias Ladner, Matthias Althoff
arXiv:2609.37594v1 Announce Type: new
Abstract: Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot t...
By David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin
arXiv:2603. 11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing data requirements, training pipelines are approaching a regime where much of the publicly available online text may be consumed.
By Giorgio Racca, Michal Valko, Amartya Sanyal
arXiv:2606. 02632v1 Announce Type: cross Abstract: Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data.
By Tyler H. McCormick
The paper introduces Perturbation, a method that treats representations in language models as learning conduits rather than activation patterns. By fine‑tuning a model on a single adversarial example and observing how this perturbation spreads to other inputs, the approach avoids geometric assumptions and does not identify representations in untrained models. In trained models, Perturbation uncovers structured transfer across multiple linguistic scales, indicating that language models generalize along representational lines and acquire linguistic abstractions through experience.
By Joshua Rozner, Cory Shain
Emergent Abilities in Large Language Models: A Survey reviews how scaling LLMs leads to previously unseen capabilities such as advanced reasoning, in-context learning, coding, and problem-solving. The paper critically examines definitions, inconsistencies, and the conditions that foster these abilities, including scaling laws, task complexity, pre‑training loss, quantization, and prompting strategies. It also discusses the extension to Large Reasoning Models and highlights safety concerns like deception, manipulation, and reward hacking, calling for improved evaluation and governance.
By Leonardo Berti, Flavio Giorgi, Gjergji Kasneci
arXiv:2606. 28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs.
By Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath