arXiv AI

Discovering Machine Correlates of Consciousness

arXiv AI
6d ago

Initial results of the Digital Consciousness Model

The Digital Consciousness Model (DCM) is introduced as a systematic, probabilistic framework for evaluating evidence of consciousness in AI systems, drawing on multiple leading theories of consciousness. Initial results indicate that, while the evidence suggests 2024 large language models (LLMs) are not conscious, this conclusion is not decisive and is weaker than evidence against consciousness in simpler AI systems.

By Derek Shiller, Laura Duffy, Arvo Mu\~noz Mor\'an, Adri\`a Moret, Chris Percy, Hayley Clatterbuck
arXiv AI
3d ago

From cacophony to hierarchy: a principled framework for assessing AI consciousness

arXiv:2609.35618v2 Announce Type: replace Abstract: The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a c...

By Shamil Chandaria, Arvo Mu\~noz Mor\'an, Fernando Rosas, Anil Seth, Henry Shevlin, Marcus Hutter, Thore Graepel, Adam Bales, Iulia Comsa, Murray Shanahan, Ruben Laukkonen, Morten Kringelbach, Chris Frith, Shane Legg
arXiv AI
6d ago

From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia

The paper proposes a five‑layer implementation architecture for the S3Q theory of consciousness, mapping its three core conditions—grounded sensorimotor situatedness, internal simulation via a world model, and structural coherence between predictions and observations—to existing computational primitives. It integrates these components into a single pipeline that processes continuous, differentiable per‑object slot vectors, includes a developmental bootstrap sequence, and yields falsifiable predictions that cannot be produced by any subset of the architecture alone. The model further suggests that a basic sense of self emerges from linking actions to outcomes, and that behavior can be categorized into hesitation, curiosity, or avoidance based on outcome unexpectedness and valence. "whyItMatters":"The architecture provides a concrete, testable framework that unifies theoretical conditions of consciousness with practical computational components, enabling empirical investigation of machine consciousness."

By Tetiana Grinberg, Katrina Schleisman, Patryk Laurent, Bogdan Udrea, Minda Myers, Brian Aufderheide, Luis El Srouji, Doyle Groves, Kevin Schmidt
arXiv AI
Sep 10

We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consciousness

The paper argues that the current debate on machine consciousness rests on the mistaken assumption that AI systems are already the kind of entities that could possess consciousness. By distinguishing between phenomenal consciousness, introspective report, and human projective introspection, it introduces the AI Consciousness Fallacy, showing that generative models can produce first‑person linguistic traces without being conscious. It then proposes Causal Liability Theory (CLT), with CLT‑I defining liability closure as a criterion for identifying a bearer of consciousness and CLT‑II suggesting that liability closure is both necessary and sufficient for minimal phenomenal subjecthood, and demonstrates experimentally that these distinctions are tractable and can separate causal bearer structure from first‑person performance.

By Afshin Khadangi
arXiv AI
2d ago

What Can Analogy Tell Us About Artificial Consciousness?

The article explores how analogy informs judgments about consciousness, especially in the context of artificial intelligence. It introduces a causal framework that distinguishes between similarities in underlying factors and similarities in observable behavior, weighting source-target similarity by causal relevance. Applying this framework to biological systems explains why analogical support weakens as causal distance from humans increases, and to AI it shows that behavioral similarity alone offers limited evidence for consciousness due to poorly established causal correspondences.

By Keith J. Holyoak, Martin M. Monti
arXiv Machine Learning
Sep 18

Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds

The paper investigates how language models can covertly encode a hidden trait—termed subliminal learning—through seemingly unrelated outputs. By systematically measuring output co‑variation, fixed output‑vector alignment, hidden‑state readability, and causal control across a range of model sizes and prompting protocols, the authors find that fixed geometry and observational readability do not reliably predict behavior, while causal timing and multi‑token measurements reveal stronger, concept‑wide effects. These distinct properties highlight that token‑level explanations are insufficient to pinpoint the mechanism behind training‑time trait transfer.

By Barath Velmurugan
arXiv AI
Aug 20

Do Large Language Models Hallucinate Electric Fata Morganas?

The paper investigates why large language models (LLMs) produce hallucinations—outputs that are fabricated, unverifiable, or contradictory to source material—and argues that these hallucinations have philosophical implications for machine consciousness. It reviews known causes such as source‑target divergence, training‑inference discrepancies, and overfitting, and presents two empirical studies: one showing that higher temperature settings in GPT models yield plausible but incorrect answers, while lower temperatures produce accurate ones; and another demonstrating that an encoder‑only model trained on encyclopedic data answers factually without embellishment, suggesting hallucinations arise from exposure to subjective, socially diverse data rather than cognitive ability. Drawing on Turing, Searle’s Chinese Room, the frame problem, and cybernetic theory, the authors contend that a model’s self‑reports of emotion or sentience fall within the definition of hallucination, implying that any future machine consciousness may remain epistemically inaccessible because it would be indistinguishable from an advanced hallucination.

By Kristina \v{S}ekrst
arXiv AI
4d ago

Which Attention Heads are like the Human Head? Not the Ones that Compute

The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.

By Christopher Pinier, Gustaw Opie{\l}ka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson
arXiv Machine Learning
Aug 27

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

The study examines whether a decodable "empathy" direction can be used as a reliable causal lever in language models. Using EPITOME-derived facets of Recognition (cognitive) and Resonance (affective) across three instruction‑tuned LLMs, the authors find that while affective steering can partially raise affective scores, cognitive steering shows inconsistent or unmeasurable effects. The results highlight that decodability does not guarantee reliable control, especially for cognitive empathy, and that measurement sensitivity must be explicitly checked.

By Haoran Jisun