Metaphor Tracer: A Theory-Informed Analysis of Hidden States
What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token position on two properties.
The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.
What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token position on two properties.
arXiv:2406. 05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process.
arXiv:2609.37497v1 Announce Type: new Abstract: Modern transformer models excel at capturing semantic relationships through sentence embeddings, yet their ability to perform pragmatic reasoning remai...
arXiv:2408.11827v2 Announce Type: replace Abstract: Understanding how language models compose meaning from linguistic input remains a central problem in interpretability research. Mechanistic studies...
arXiv:2511. 21731v2 Announce Type: replace-cross Abstract: We present the results of cognitive tests on conceptual combinations, performed using specific Large Language Models (LLMs) as test subjects.
arXiv:2603. 20381v2 Announce Type: replace-cross Abstract: Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designing safe, thoughtful, engaging, and empowering human-agent interactions.
The study investigates whether pretrained transformer models encode functional words—such as pronouns and adverbs—in a way that mirrors human usage. By comparing embeddings of nouns with those of their functional counterparts in both isolated and parallel sentences, the authors find that functional words occupy a central yet distinct position in embedding space and that parallel lexicalized and functional sentences reside in different subspaces. Experiments show that only a mixed training set of functional and lexicalized sentences reveals shared syntactic and semantic structure, whereas training on either type alone fails to capture this parallelism.
arXiv:2608. 11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways.
arXiv:2607. 07891v1 Announce Type: cross Abstract: Roy Harris's Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computational approaches to language, arguing that language is not a code that maps onto a pre-given world but a situated, bipartite activity oriented toward prospective joint action.
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
The study investigates how disrupting conceptual versus referential information in short narratives affects human reading and large language model (LLM) processing. In humans, conceptual disruptions cause a strong, localized processing cost that peaks early and declines quickly, while referential disruptions produce weaker, gradually decreasing effects that are more influenced by sentence boundaries. In LLMs, both disruptions appear immediately at the manipulated word; surprisal patterns mirror human reading, whereas output-layer representations show that referential disruption initially causes a larger displacement before both types decay following a power-law.
arXiv:2609.07474v2 Announce Type: replace Abstract: Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should...