arXiv AI

Signatures of Steerability in Activation Space of Language Models

arXiv AI
Jun 2

Concept Heterogeneity-aware Representation Steering

arXiv:2603. 02237v2 Announce Type: replace-cross Abstract: Representation steering offers a lightweight mechanism for controlling the behavior of large language models (LLMs) by intervening on internal activations at inference time.

By Laziz U. Abdullaev, Noelle Y. L. Wong, Ryan T. Z. Lee, Shiqi Jiang, Khoi N. M. Nguyen, Tan M. Nguyen
arXiv Machine Learning
Jul 9

Towards Understanding Steering Strength

arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.

By Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
arXiv Machine Learning
Sep 10

Disentangling Steering Vectors

arXiv:2609.07037v1 Announce Type: new Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...

By Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi, Hisashi Kashima
arXiv AI
Sep 2

Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures

The paper investigates whether neural networks exhibit conceptual separation, meaning that examples of the same concept cluster together and related concepts are closer in representation space. Using geometric and distributional analyses, the authors find that Convolutional Neural Networks (CNNs) produce coherent, semantically ordered representations for familiar ImageNet concepts, but this coherence weakens for unseen concepts and under domain shift. Large Language Models (LLMs) keep distinct domains well separated, bring related subdomains closer, yet lose distinction between ambiguous topics at both mean and covariance levels.

By Jaee Ponde, Roshni Agarwal, Subhashis Banerjee