Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
arXiv:2608. 07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization.
MAxBench is a geometry‑agnostic benchmark for evaluating how well language models recover multinomial concept representations. The study compares ten localization methods across five geometry types, six concepts, and four models, finding that affine subspaces generally steer more reliably and recall more instances than rank‑one or linear subspaces. The results also show that manifold steering can match the best methods when applicable, and that no method consistently outperforms prompting for these complex concepts.
arXiv:2608. 07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization.
arXiv:2609.38625v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remai...
The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.
arXiv:2608.31084v1 Announce Type: new Abstract: The Jacobian Lens (J-lens) is a recent tool for interpreting LLMs. It reads a hidden state as a ranked list of vocabulary tokens, leaving multi-token c...
arXiv:2512. 07355v2 Announce Type: replace Abstract: Two traditions of interpretability have evolved side by side but seldom spoken to each other: Concept Bottleneck Models (CBMs), which prescribe what a concept should be, and Sparse Autoencoders (SAEs), which discover what concepts emerge.
arXiv:2606. 09653v1 Announce Type: new Abstract: Learned representations across models and modalities often exhibit striking structural similarities, suggesting shared underlying concept decompositions.
arXiv:2607. 10578v1 Announce Type: new Abstract: Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge.
arXiv:2607. 04222v1 Announce Type: new Abstract: Interpretability methods aim to reveal the features represented inside large language models (LLMs).
arXiv:2607. 08284v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them.
arXiv:2608. 20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely.
arXiv:2606. 11722v1 Announce Type: cross Abstract: Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior.
arXiv:2606. 11105v1 Announce Type: cross Abstract: Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them.