The Digital Consciousness Model (DCM) is introduced as a systematic, probabilistic framework for evaluating evidence of consciousness in AI systems, drawing on multiple leading theories of consciousness. Initial results indicate that, while the evidence suggests 2024 large language models (LLMs) are not conscious, this conclusion is not decisive and is weaker than evidence against consciousness in simpler AI systems.
By Derek Shiller, Laura Duffy, Arvo Mu\~noz Mor\'an, Adri\`a Moret, Chris Percy, Hayley Clatterbuck
arXiv:2603.27597v2 Announce Type: replace
Abstract: Research on artificial consciousness increasingly shifts evaluation from behaviour to internal architecture. Theory-based indicators are used to up...
By Florentin Koch
arXiv:2609.35618v2 Announce Type: replace
Abstract: The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a c...
By Shamil Chandaria, Arvo Mu\~noz Mor\'an, Fernando Rosas, Anil Seth, Henry Shevlin, Marcus Hutter, Thore Graepel, Adam Bales, Iulia Comsa, Murray Shanahan, Ruben Laukkonen, Morten Kringelbach, Chris Frith, Shane Legg
The paper proposes a five‑layer implementation architecture for the S3Q theory of consciousness, mapping its three core conditions—grounded sensorimotor situatedness, internal simulation via a world model, and structural coherence between predictions and observations—to existing computational primitives. It integrates these components into a single pipeline that processes continuous, differentiable per‑object slot vectors, includes a developmental bootstrap sequence, and yields falsifiable predictions that cannot be produced by any subset of the architecture alone. The model further suggests that a basic sense of self emerges from linking actions to outcomes, and that behavior can be categorized into hesitation, curiosity, or avoidance based on outcome unexpectedness and valence.
"whyItMatters":"The architecture provides a concrete, testable framework that unifies theoretical conditions of consciousness with practical computational components, enabling empirical investigation of machine consciousness."
By Tetiana Grinberg, Katrina Schleisman, Patryk Laurent, Bogdan Udrea, Minda Myers, Brian Aufderheide, Luis El Srouji, Doyle Groves, Kevin Schmidt
The paper argues that the current debate on machine consciousness rests on the mistaken assumption that AI systems are already the kind of entities that could possess consciousness. By distinguishing between phenomenal consciousness, introspective report, and human projective introspection, it introduces the AI Consciousness Fallacy, showing that generative models can produce first‑person linguistic traces without being conscious. It then proposes Causal Liability Theory (CLT), with CLT‑I defining liability closure as a criterion for identifying a bearer of consciousness and CLT‑II suggesting that liability closure is both necessary and sufficient for minimal phenomenal subjecthood, and demonstrates experimentally that these distinctions are tractable and can separate causal bearer structure from first‑person performance.
By Afshin Khadangi
arXiv:2608. 19215v1 Announce Type: new Abstract: Given deep uncertainty about the possibility of artificial consciousness, it is unclear how we should treat potentially sentient AI.
By Dr Tom McClelland
The article explores how analogy informs judgments about consciousness, especially in the context of artificial intelligence. It introduces a causal framework that distinguishes between similarities in underlying factors and similarities in observable behavior, weighting source-target similarity by causal relevance. Applying this framework to biological systems explains why analogical support weakens as causal distance from humans increases, and to AI it shows that behavioral similarity alone offers limited evidence for consciousness due to poorly established causal correspondences.
By Keith J. Holyoak, Martin M. Monti
arXiv:2601. 15334v2 Announce Type: replace-cross Abstract: Whether language models possess sentience has no empirical answer.
By Caspar Kaiser, Sean Enderby
The paper investigates how language models can covertly encode a hidden trait—termed subliminal learning—through seemingly unrelated outputs. By systematically measuring output co‑variation, fixed output‑vector alignment, hidden‑state readability, and causal control across a range of model sizes and prompting protocols, the authors find that fixed geometry and observational readability do not reliably predict behavior, while causal timing and multi‑token measurements reveal stronger, concept‑wide effects. These distinct properties highlight that token‑level explanations are insufficient to pinpoint the mechanism behind training‑time trait transfer.
By Barath Velmurugan
The paper investigates why large language models (LLMs) produce hallucinations—outputs that are fabricated, unverifiable, or contradictory to source material—and argues that these hallucinations have philosophical implications for machine consciousness. It reviews known causes such as source‑target divergence, training‑inference discrepancies, and overfitting, and presents two empirical studies: one showing that higher temperature settings in GPT models yield plausible but incorrect answers, while lower temperatures produce accurate ones; and another demonstrating that an encoder‑only model trained on encyclopedic data answers factually without embellishment, suggesting hallucinations arise from exposure to subjective, socially diverse data rather than cognitive ability. Drawing on Turing, Searle’s Chinese Room, the frame problem, and cybernetic theory, the authors contend that a model’s self‑reports of emotion or sentience fall within the definition of hallucination, implying that any future machine consciousness may remain epistemically inaccessible because it would be indistinguishable from an advanced hallucination.
By Kristina \v{S}ekrst
The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.
By Christopher Pinier, Gustaw Opie{\l}ka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson
The study examines whether a decodable "empathy" direction can be used as a reliable causal lever in language models. Using EPITOME-derived facets of Recognition (cognitive) and Resonance (affective) across three instruction‑tuned LLMs, the authors find that while affective steering can partially raise affective scores, cognitive steering shows inconsistent or unmeasurable effects. The results highlight that decodability does not guarantee reliable control, especially for cognitive empathy, and that measurement sensitivity must be explicitly checked.
By Haoran Jisun