arXiv:2609.24821v1 Announce Type: new
Abstract: The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-relate...
By Manjiang Yu, Hongji Li, Zihan Wang, Junwei Chen, Xue Li, Priyanka Singh, Yang Cao, Lijie Hu
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
By Max Torop, Aria Masoomi, Jennifer Dy
How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects.
arXiv:2607. 04525v1 Announce Type: cross Abstract: How concepts are represented in neural networks is a fundamental question in machine learning.
By Zhimin Hu, Lanhao Niu, Sashank Varma
arXiv:2609.16854v1 Announce Type: new
Abstract: Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing...
By Ferm\'{\i}n Moscoso del Prado Mart\'{\i}n
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty...
The paper introduces a method that represents language models as log‑likelihood vectors over prompt‑response pairs, enabling the construction of model maps that compare conditional distributions. Squared Euclidean distances in this vector space approximate KL divergence, and experiments show that these maps reveal global structure related to model attributes and task performance. The approach also captures systematic shifts from prompt changes, supports additive compositionality for predicting downstream scores, and offers PMI vectors to mitigate unconditional distribution effects, thereby aiding analysis and prediction of input‑dependent behavior.
By Momose Oyama, Yusuke Takase, Hidetoshi Shimodaira
arXiv:2609.34187v2 Announce Type: replace-cross
Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
By Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.
By Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.
By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
The paper investigates how large language models encode and use relational information among tokens across transformer layers. By analyzing activations from prompts that require inferring relationships among three cyclic tokens (months, hours, weekdays, musical notes), the authors find a consistent layerwise progression: intermediate layers capture pairwise relationships, while later layers encode the full three‑token relationship to predict the next token. They also identify geometrically structured token relationships that do not influence prediction, and show that constraining models to use only causally relevant joint representations improves next‑token accuracy.
By Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls