arXiv:2603.18007v2 Announce Type: replace-cross
Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to inf...
By Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
By Anthony Baez, Sheer Karny, Pat Pataranutaporn
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
The paper introduces a new method for measuring metacognitive abilities in large language models (LLMs) without relying on self-reports, instead testing how well models can use knowledge of their internal states. Using two experimental paradigms, the authors find that recent frontier LLMs can assess and use their own confidence when answering factual and reasoning questions, and can anticipate and appropriately employ the answers they would give. The study also shows that these abilities are limited in resolution, context-dependent, differ qualitatively from human metacognition, and vary across models with similar capabilities, suggesting post‑training processes influence metacognitive development.
By Christopher Ackerman
arXiv:2606. 32008v1 Announce Type: new Abstract: Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.
By Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo, Chia-Tse Shao, Yingxiao Ye, Aobo Yang, Vivek Miglani, Nehal Bandi
arXiv:2609.07943v1 Announce Type: new
Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...
By Alex Smolin, Bryan Wilder
arXiv:2607. 09306v2 Announce Type: replace-cross Abstract: Whether a language model behaves as it claims is a judgement on which independent human raters cannot agree (Fleiss kappa = 0.
By Kwan Soo Shin, In Seok Kang, Yunkyung Min, Munho Lee
The paper investigates why large language models (LLMs) produce hallucinations—outputs that are fabricated, unverifiable, or contradictory to source material—and argues that these hallucinations have philosophical implications for machine consciousness. It reviews known causes such as source‑target divergence, training‑inference discrepancies, and overfitting, and presents two empirical studies: one showing that higher temperature settings in GPT models yield plausible but incorrect answers, while lower temperatures produce accurate ones; and another demonstrating that an encoder‑only model trained on encyclopedic data answers factually without embellishment, suggesting hallucinations arise from exposure to subjective, socially diverse data rather than cognitive ability. Drawing on Turing, Searle’s Chinese Room, the frame problem, and cybernetic theory, the authors contend that a model’s self‑reports of emotion or sentience fall within the definition of hallucination, implying that any future machine consciousness may remain epistemically inaccessible because it would be indistinguishable from an advanced hallucination.
By Kristina \v{S}ekrst
arXiv:2606. 12618v1 Announce Type: new Abstract: Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say.
By Alan Cooney, David Africa, Geoffrey Irving
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief.
arXiv:2608.17809v2 Announce Type: replace
Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevita...
By Quang Minh Nguyen, Luis Frentzen Salim
arXiv:2606. 06380v1 Announce Type: cross Abstract: The question of whether artificial systems can be conscious remains open, in part because existing approaches either evaluate systems against theory-derived checklists (discriminative) or engineer consciousness-inspired modules directly (architectural); both leave open whether observed structures are artifacts of human language priors.
By Zengqing Wu, Chuan Xiao