arXiv AI
Sep 2

In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?

The paper investigates whether large language models (LLMs) can control their own internal representations, a question relevant to machine metacognition and AI safety. Previous work claimed such control using neurofeedback, but the targets were not privileged, allowing third‑party inference from prompts. By redesigning the paradigm to enforce privileged access—mirroring human neurofeedback experiments—the authors find that LLMs fail to reliably control these internal representations, indicating that earlier claims may have relied on superficial mechanisms.

By Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi, Yusuke Haruki, Daisuke Kawahara
arXiv AI
Aug 24

Can LLMs Introspect? A Reality Check

The paper questions whether large language models (LLMs) truly introspect by critiquing recent studies that claim they can detect and report their internal states. It proposes two necessary conditions for genuine introspection: privileged access to internal representations and second‑order computation that distinguishes from first‑order task performance. Re‑examining two existing paradigms, the authors find that apparent introspective abilities can be explained by input‑based classifiers or generic anomaly detection, concluding that current evidence does not support metacognitive monitoring in LLMs.

By Shashwat Singh, Tal Linzen, Shauli Ravfogel
arXiv Machine Learning
Sep 11

Evidence for Limited Metacognition in LLMs

The paper introduces a new method for measuring metacognitive abilities in large language models (LLMs) without relying on self-reports, instead testing how well models can use knowledge of their internal states. Using two experimental paradigms, the authors find that recent frontier LLMs can assess and use their own confidence when answering factual and reasoning questions, and can anticipate and appropriately employ the answers they would give. The study also shows that these abilities are limited in resolution, context-dependent, differ qualitatively from human metacognition, and vary across models with similar capabilities, suggesting post‑training processes influence metacognitive development.

By Christopher Ackerman
arXiv AI
Jun 16

Metacognitive Myopia in Large Language Models

arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.

By Florian Scholten, Tobias R. Rebholz, Mandy H\"utter