The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
arXiv:2601. 06599v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations.
arXiv:2607. 16741v1 Announce Type: new Abstract: B\"urger et al.
The paper investigates how truth representations in small language models are structured. Using a training‑free axis derived from the dominant singular vector of hidden‑state differences between true and false minimal pairs, the authors evaluate 14 models across six architectural families, including Mixture‑of‑Experts. The study examines whether a single direction captures truth, which components contribute, and how this applies to categories with computed truth values.
arXiv:2606. 15821v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, forming distinct model lineages.
arXiv:2604.03754v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous...
arXiv:2608. 06417v1 Announce Type: new Abstract: The proliferation of misinformation online has driven demand for scalable detection systems.
arXiv:2610.00910v1 Announce Type: cross Abstract: Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual...
arXiv:2606. 03022v1 Announce Type: cross Abstract: Hallucination in Large Language Models (LLMs), characterized by the generation of content inconsistent with contextual facts or logical constraints -- remains a persistent challenge for reliable deployment.
FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
arXiv:2606. 24964v1 Announce Type: new Abstract: Understanding the features of large language models (LLMs) is a central goal of interpretability.
The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between mo...
Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed world-model the network stores and looks up. We show that context is king: the structure a model actually uses is set by the in-context specification.