One mechanism for many mental spaces: a shared router over a value slot in language models
arXiv:2607. 10248v1 Announce Type: cross Abstract: Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical.
arXiv:2607. 11945v1 Announce Type: cross Abstract: Capable language models hold what a character believes apart from what is true: told "Anna believes the cup is blue; in reality it is red," they answer blue about Anna and red about the world.
arXiv:2607. 10248v1 Announce Type: cross Abstract: Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical.
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
arXiv:2607. 15495v1 Announce Type: cross Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning.
arXiv:2610.00910v1 Announce Type: cross Abstract: Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual...
The study turns the speculative field of astrolinguistics into an experiment by using two large language models with deliberately incompatible constructed lexicons as informants. A scripted orchestrator translates between the two category systems, and a protocol combining cross‑situational elimination, predictive probes, active scene selection, and a stricter recovery round successfully prevents the ‘kangaroo effect’—the silent attachment of a word to the wrong referent—in over 400 simulated and live runs. When informant noise is introduced, the protocol remains robust up to 2% per‑word noise and largely abstains rather than errs at higher noise levels, while a generate‑and‑test loop allows recovery of words outside the scripted hypothesis space, achieving full coverage as the rule‑proposing LLM’s capability increases. whyItMatters":"The protocol demonstrates that experimental astrolinguistics can reliably avoid mistranslations and recover unknown terms, showing that correctness is governed by the protocol while coverage depends on the instruments used."
arXiv:2607. 15883v1 Announce Type: cross Abstract: Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind.
The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between mo...
The paper investigates how truth representations in small language models are structured. Using a training‑free axis derived from the dominant singular vector of hidden‑state differences between true and false minimal pairs, the authors evaluate 14 models across six architectural families, including Mixture‑of‑Experts. The study examines whether a single direction captures truth, which components contribute, and how this applies to categories with computed truth values.
The study investigates how demographic identity is represented in a language model, using representational similarity analysis against Pew survey data across 169 demographic cells. It finds that standard last‑token read‑outs underestimate the model’s fidelity, while specific attention heads (notably L11 H16) capture demographic structure more accurately, though race‑based types remain weak. Causal interventions reveal that high fidelity does not guarantee causal use, and a 128‑dimensional probe of a single head improves alignment with survey truth but fails to recover per‑question group ordering.
In long, multi-turn dialogue a large language model maintains an implicit relational stance toward the user, spanning from "push the user toward real-world others" to "position itself as the user's sole support. " When it slides toward the latter, "support" degrades into "you only have me" -- a harm documented in real companion conversations (Moore et al.
arXiv:2607. 03598v1 Announce Type: cross Abstract: When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was doing by sending it: share a finished project and it critiques the code; share a raw late-night line and it runs a wellness check.
arXiv:2609.24209v1 Announce Type: new Abstract: The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent wor...