Inverted Detection and Control in Steering Vectors
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
arXiv:2607. 04525v1 Announce Type: cross Abstract: How concepts are represented in neural networks is a fundamental question in machine learning.
How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects.
arXiv:2609.16854v1 Announce Type: new Abstract: Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing...
arXiv:2607. 20454v1 Announce Type: cross Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation.
arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty...
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2606. 15733v1 Announce Type: cross Abstract: Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer are unchanged.
The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.
The paper introduces a method that represents language models as log‑likelihood vectors over prompt‑response pairs, enabling the construction of model maps that compare conditional distributions. Squared Euclidean distances in this vector space approximate KL divergence, and experiments show that these maps reveal global structure related to model attributes and task performance. The approach also captures systematic shifts from prompt changes, supports additive compositionality for predicting downstream scores, and offers PMI vectors to mitigate unconditional distribution effects, thereby aiding analysis and prediction of input‑dependent behavior.
arXiv:2608. 07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization.