The paper investigates the relationship between a large language model’s internal probability distribution and its verbalized confidence statements. By systematically manipulating training and in‑context data, the authors show that both internal and verbalized probabilities are influenced by distributional and asserted uncertainty in the data. They find that verbalized probabilities align with internal ones beyond what would be expected if they tracked the same sources independently, indicating that verbalized confidence can serve as a probe of the model’s internal distribution.
By Sinead Williamson, Jiaxuan Li, Nick Foti, Russ Webb, Masha Fedzechkina
arXiv:2603.18908v5 Announce Type: replace
Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
By Matt Gorbett, Suman Jana
arXiv:2506. 14003v5 Announce Type: replace Abstract: Machine unlearning (MU) for large language models (LLMs), commonly referred to as LLM unlearning, seeks to remove specific undesirable data or knowledge from a trained model, while maintaining its performance on standard tasks.
By Yiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu, Sijia Liu
arXiv:2606. 07951v1 Announce Type: cross Abstract: Humans increasingly turn to Language Models (LMs) in ways that shape beliefs and drive decisions, including discussing, rewriting, and summarizing information from scientific articles, news, and medical reports.
By Catarina G Belem, Shang Wu, Hongyu Yao, Mark Steyvers, Sameer Singh, Padhraic Smyth
arXiv:2609.34187v2 Announce Type: replace-cross
Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
By Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
arXiv:2609.15533v1 Announce Type: cross
Abstract: Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inheren...
By Tobias Ladner, Matthias Althoff