The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.
arXiv:2510. 04120v2 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing.
By Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong
arXiv:2408.11827v2 Announce Type: replace
Abstract: Understanding how language models compose meaning from linguistic input remains a central problem in interpretability research. Mechanistic studies...
By Nura Aljaafari, Danilo S. Carvalho, Andr\'e Freitas
The paper introduces Contrastive Projection, a method that reads a transformer’s internal states by differencing the hidden states of two closely matched prompts and projecting the difference through the unembedding layer. This approach cancels shared components and highlights the distinctions between prompts, effectively revealing steering vectors and domain-to-domain mappings such as metaphor. The technique is training‑free, operates at every position, sub‑layer, and head, and has been validated across multiple architectures and initialization seeds.
By Olli Tuomi
The study investigates how disrupting conceptual versus referential information in short narratives affects human reading and large language model (LLM) processing. In humans, conceptual disruptions cause a strong, localized processing cost that peaks early and declines quickly, while referential disruptions produce weaker, gradually decreasing effects that are more influenced by sentence boundaries. In LLMs, both disruptions appear immediately at the manipulated word; surprisal patterns mirror human reading, whereas output-layer representations show that referential disruption initially causes a larger displacement before both types decay following a power-law.
By Rui He, Nihal Altay, Wolfram Hinzen
MechaTerp-TRACE is a new framework that systematically ablates individual components of language models to measure their causal contribution to producing a named entity. By applying TRACE to thirteen instruction‑tuned dense decoder models, the study finds that a small set of positionally fixed components consistently carry the most influence across models and prompts, while the remaining support is evenly distributed. This suggests that entity knowledge is largely embedded in generic generation machinery rather than in isolated, findable components.
By Brandon Colelough, Davis Bartels, Madeline Bittner, Dina Demner-Fushman
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
By Francesco Karim Vicidomini
The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.
By Arman Nik Khah
arXiv:2607. 10248v1 Announce Type: cross Abstract: Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical.
By Oliver Steele, Jiangtao Wen, Yuxing Han
arXiv:2609.07474v2 Announce Type: replace
Abstract: Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should...
By Peng Xie, Amr Alanwar
The study investigates the internal workings of an audio language model (Qwen3-Omni) by applying a logit lens to its middle layers. It finds that the model’s reasoning about spoken questions becomes legible in words before any token is emitted, revealing language‑agnostic, paralinguistic, and temporally distinct signals that are causally used in the network’s decision process. The authors demonstrate that these signals can be isolated and mapped to specific layers, providing a qualitative account of how the model processes audio input.
By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Qi Luo, Jia-Hong Huang, M. Maruf, Roger Ren, Yile Gu, Rahul Pandey, Ge Liu, Ivan Bulyko
arXiv:2608. 04021v1 Announce Type: cross Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless of where the readout slot sits.
By Han-yu Wang