Hugging Face Trending Papers

Metaphor Tracer: A Theory-Informed Analysis of Hidden States

What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token position on two properties.

Hugging Face Trending Papers
Aug 18

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.

arXiv Computation and Language
Sep 10

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

The paper introduces Contrastive Projection, a method that reads a transformer’s internal states by differencing the hidden states of two closely matched prompts and projecting the difference through the unembedding layer. This approach cancels shared components and highlights the distinctions between prompts, effectively revealing steering vectors and domain-to-domain mappings such as metaphor. The technique is training‑free, operates at every position, sub‑layer, and head, and has been validated across multiple architectures and initialization seeds.

By Olli Tuomi
arXiv Computation and Language
Aug 27

Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing

The study investigates how disrupting conceptual versus referential information in short narratives affects human reading and large language model (LLM) processing. In humans, conceptual disruptions cause a strong, localized processing cost that peaks early and declines quickly, while referential disruptions produce weaker, gradually decreasing effects that are more influenced by sentence boundaries. In LLMs, both disruptions appear immediately at the manipulated word; surprisal patterns mirror human reading, whereas output-layer representations show that referential disruption initially causes a larger displacement before both types decay following a power-law.

By Rui He, Nihal Altay, Wolfram Hinzen
arXiv Computation and Language
Sep 22

MechaTerp-TRACE: A Novel Approach for Component Ablation Analysis in Language Models

MechaTerp-TRACE is a new framework that systematically ablates individual components of language models to measure their causal contribution to producing a named entity. By applying TRACE to thirteen instruction‑tuned dense decoder models, the study finds that a small set of positionally fixed components consistently carry the most influence across models and prompts, while the remaining support is evenly distributed. This suggests that entity knowledge is largely embedded in generic generation machinery rather than in isolated, findable components.

By Brandon Colelough, Davis Bartels, Madeline Bittner, Dina Demner-Fushman
arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah
arXiv Computation and Language
Aug 27

Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

The study investigates the internal workings of an audio language model (Qwen3-Omni) by applying a logit lens to its middle layers. It finds that the model’s reasoning about spoken questions becomes legible in words before any token is emitted, revealing language‑agnostic, paralinguistic, and temporally distinct signals that are causally used in the network’s decision process. The authors demonstrate that these signals can be isolated and mapped to specific layers, providing a qualitative account of how the model processes audio input.

By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Qi Luo, Jia-Hong Huang, M. Maruf, Roger Ren, Yile Gu, Rahul Pandey, Ge Liu, Ivan Bulyko