arXiv:2607. 11881v1 Announce Type: cross Abstract: Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more.
By Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, Jordan Thomas, Mark Steyvers, Arman Cohan
arXiv:2607. 29093v1 Announce Type: cross Abstract: Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement.
By Saurabh Ranjan, Mukesh Makwana, Konstantina Sokratous, Brian Odegaard
arXiv:2608. 15400v1 Announce Type: new Abstract: Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizing when problems exceed their expertise; such limitations inevitably undermine reliability and trust in LLMs.
By Charles Courchaine, Ricky J. Sethi, Hefei Qiu
arXiv:2606. 32032v1 Announce Type: cross Abstract: Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes.
By Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona, Idan Szpektor, Arman Cohan
arXiv:2603. 25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity).
By Jon-Paul Cacioli
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent?
arXiv:2607. 12447v1 Announce Type: cross Abstract: Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models.
By Dharshan Kumaran, Viorica Patraucean, Maks Ovsanikov, Petar Veli\v{c}kovi\'c, Nathaniel Daw
arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.
By Florian Scholten, Tobias R. Rebholz, Mandy H\"utter
arXiv:2607. 20065v1 Announce Type: new Abstract: Enterprise strategic decision support requires AI systems that are not only accurate, but also uncertainty-aware, risk-calibrated, explainable, and governance-compliant.
By Tian Qiu, Li Yan, Mahabubur Rahman Miraj, Shanqin Yi, Md Intekhab Rahman Galib, Jahid Hasan
arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.
By Jinseok Chung, Minkyoung Song, Hyunji Jung, Namhoon Lee
arXiv:2607. 04034v1 Announce Type: cross Abstract: The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks.
By Stefan Broecker, Mason del Rosario, Boris Selitser, Thomas Strohmer