Disentangling Steering Vectors
arXiv:2609.07037v1 Announce Type: new Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...
arXiv:2512. 15134v2 Announce Type: replace-cross Abstract: A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks.
arXiv:2609.07037v1 Announce Type: new Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...
arXiv:2606. 29888v1 Announce Type: new Abstract: Vision-language models map images and text into a joint embedding space.
arXiv:2606. 05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information.
To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information. Learning disentangled representations is a principled way to identify these underlying shared and unique factors that are hidden in observational data.
arXiv:2607. 17770v1 Announce Type: cross Abstract: Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations.
arXiv:2601. 21944v3 Announce Type: replace Abstract: The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability.
The paper investigates whether neural networks exhibit conceptual separation, meaning that examples of the same concept cluster together and related concepts are closer in representation space. Using geometric and distributional analyses, the authors find that Convolutional Neural Networks (CNNs) produce coherent, semantically ordered representations for familiar ImageNet concepts, but this coherence weakens for unseen concepts and under domain shift. Large Language Models (LLMs) keep distinct domains well separated, bring related subdomains closer, yet lose distinction between ambiguous topics at both mean and covariance levels.
arXiv:2509. 22015v2 Announce Type: replace Abstract: Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, providing a powerful lens for passive feature discovery.
arXiv:2608. 10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc.
arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.
arXiv:2608.30563v1 Announce Type: new Abstract: Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios. Existing methods usually generate, align, or...
arXiv:2607. 00023v1 Announce Type: cross Abstract: Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to feature superposition.