arXiv Computation and Language By Olli Tuomi

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

Read the original on arXiv Computation and Language →

The paper introduces Contrastive Projection, a method that reads a transformer’s internal states by differencing the hidden states of two closely matched prompts and projecting the difference through the unembedding layer. This approach cancels shared components and highlights the distinctions between prompts, effectively revealing steering vectors and domain-to-domain mappings such as metaphor. The technique is training‑free, operates at every position, sub‑layer, and head, and has been validated across multiple architectures and initialization seeds.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 3

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

The paper introduces Sparse Readout Prism (SRP), a method that decomposes a language model’s readout matrix into sparse features, allowing logit‑lens scores to be expressed as sums of feature contributions. SRP reveals that lens readings depend on the corpus used to fit the readout, a phenomenon called corpus conditionality, and that the dominant readout feature remains stable across different corpora. By replacing the original readout with SRP’s sparse approximation, the authors recover 8.9–17.3 percentage points more of the tested logit differences than six geometric‑relation baselines, and ablating features shifts logit differences proportionally to their SRP contributions.

By Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane