arXiv AI By Elan Barenholtz

World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Sep 22

Replicating the Geometry of Emotion Representations in a Base Open-Weights Model

Sofroniew et al. (2026) showed that emotion concepts in Claude Sonnet 4.5 are encoded as vectors whose geometry mirrors human affect psychology. This study replicates that finding using the base pretrained model google/gemma-2-27b, generating 205,200 Claude Sonnet 4.5 stories, extracting 171 emotion vectors, and recovering a similar affective circumplex with principal components explaining comparable variance. The analysis further identifies a sharp geometric seam at layers 22‑26, demonstrates that much of the geometry already exists in static token embeddings, and shows that the geometry predicts token‑level co‑activation with high correlation.

By Adam Hollowell
arXiv Computation and Language
Sep 18

Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

The study investigates whether pretrained transformer models encode functional words—such as pronouns and adverbs—in a way that mirrors human usage. By comparing embeddings of nouns with those of their functional counterparts in both isolated and parallel sentences, the authors find that functional words occupy a central yet distinct position in embedding space and that parallel lexicalized and functional sentences reside in different subspaces. Experiments show that only a mixed training set of functional and lexicalized sentences reveals shared syntactic and semantic structure, whereas training on either type alone fails to capture this parallelism.

By Giuseppe Samo, Vivi Nastase, Paola Merlo
arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte
arXiv AI
Sep 7

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

The article presents a technical manual for an open toolkit designed to measure how transformer language models individuate word meanings across different contexts. It introduces the concept of a "bridge form"—a single word that appears unchanged in multiple domains but with distinct senses—and outlines a full pipeline from specifying these forms to extracting layer-wise representations, computing silhouette-based separation metrics, and visualizing results. The manual details each design choice and its intended methodological safeguards, emphasizing that it serves as a methodological reference rather than reporting empirical findings.

By Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Marcelo Vinicius de Paula, T\'arcio Andr\'e dos Santos Barros