arXiv AI

Metacognition in LLMs: Foundations, Progress, and Opportunities

arXiv:2607. 11881v1 Announce Type: cross Abstract: Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more.

arXiv Machine Learning
Sep 11

Evidence for Limited Metacognition in LLMs

The paper introduces a new method for measuring metacognitive abilities in large language models (LLMs) without relying on self-reports, instead testing how well models can use knowledge of their internal states. Using two experimental paradigms, the authors find that recent frontier LLMs can assess and use their own confidence when answering factual and reasoning questions, and can anticipate and appropriately employ the answers they would give. The study also shows that these abilities are limited in resolution, context-dependent, differ qualitatively from human metacognition, and vary across models with similar capabilities, suggesting post‑training processes influence metacognitive development.

By Christopher Ackerman
arXiv AI
Jul 9

Measuring the metacognition of AI

arXiv:2603. 29693v3 Announce Type: replace Abstract: A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks.

By Richard Servajean, Philippe Servajean
arXiv Computation and Language
4d ago

LLMs learn different forms of metacognition when trained to predict their own accuracy

The study trains ten open‑weight large language models (LLMs) to predict their own accuracy on factual multiple‑choice questions before answering. Results show that the models’ confidence signals split into two distinct patterns: early in training, confidence aligns with output consistency (how concentrated the answer distribution is), while later, it aligns with true accuracy but only on data similar to the training set. This indicates that calibration training may not universally teach LLMs to detect their own errors.

By Nicolas Yax, Stefano Palminteri, Pierre-Yves Oudeyer
arXiv Machine Learning
1d ago

Metacognitive Reasoning in Energy Based Models using Instance Based Learning Theory

The paper introduces MERITED, a framework that combines Instance-Based Learning Theory (IBLT) with Energy Based Models (EBMs) to enable metacognitive reasoning about computational effort. It allows an EBM to dynamically allocate resources based on uncertainty estimates, addressing limitations of large language models that cannot predict uncertainty before responding. The authors present a 191‑million‑parameter reasoning EBM and demonstrate how MERITED uses an IBL model for efficient, uncertainty‑driven compute allocation.

By Tailia Malloy, Prateek Kumar Rajput, Serge Lionel Nikiema, Cleotilde Gonzalez, Tegawend\'e F. Bissyand\'e
arXiv AI
Aug 24

Can LLMs Introspect? A Reality Check

The paper questions whether large language models (LLMs) truly introspect by critiquing recent studies that claim they can detect and report their internal states. It proposes two necessary conditions for genuine introspection: privileged access to internal representations and second‑order computation that distinguishes from first‑order task performance. Re‑examining two existing paradigms, the authors find that apparent introspective abilities can be explained by input‑based classifiers or generic anomaly detection, concluding that current evidence does not support metacognitive monitoring in LLMs.

By Shashwat Singh, Tal Linzen, Shauli Ravfogel