metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making
arXiv:2607. 29093v1 Announce Type: cross Abstract: Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement.
arXiv:2603. 25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity).
arXiv:2607. 29093v1 Announce Type: cross Abstract: Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement.
arXiv:2606. 32032v1 Announce Type: cross Abstract: Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes.
arXiv:2603.17839v4 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...
arXiv:2608. 14552v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty.
The paper introduces a new method for measuring metacognitive abilities in large language models (LLMs) without relying on self-reports, instead testing how well models can use knowledge of their internal states. Using two experimental paradigms, the authors find that recent frontier LLMs can assess and use their own confidence when answering factual and reasoning questions, and can anticipate and appropriately employ the answers they would give. The study also shows that these abilities are limited in resolution, context-dependent, differ qualitatively from human metacognition, and vary across models with similar capabilities, suggesting post‑training processes influence metacognitive development.
arXiv:2606. 19509v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own knowledge on such tasks remains unexplored.
arXiv:2606. 28881v1 Announce Type: cross Abstract: Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring systems.
arXiv:2608.03854v4 Announce Type: replace Abstract: Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability...
The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.
arXiv:2603. 29693v3 Announce Type: replace Abstract: A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks.
arXiv:2605. 27752v2 Announce Type: replace Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence.
The study trains ten open‑weight large language models (LLMs) to predict their own accuracy on factual multiple‑choice questions before answering. Results show that the models’ confidence signals split into two distinct patterns: early in training, confidence aligns with output consistency (how concentrated the answer distribution is), while later, it aligns with true accuracy but only on data similar to the training set. This indicates that calibration training may not universally teach LLMs to detect their own errors.