High-Dimensional Random Projection for Activation Steering in Language Models
arXiv:2606. 15092v1 Announce Type: new Abstract: Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs).
arXiv:2602. 00945v2 Announce Type: replace-cross Abstract: LLMs are multilingual by training, yet their lingua franca is often English, reflecting English language dominance in pretraining.
arXiv:2606. 15092v1 Announce Type: new Abstract: Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs).
arXiv:2608. 12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.
arXiv:2604. 03532v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult.
FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
arXiv:2609.31122v1 Announce Type: new Abstract: Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate...
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
The paper introduces a new method for identifying language-specific neurons in multilingual large language models (mLLMs). Unlike previous entropy-based approaches that only consider positive activations, the proposed Distribution-aware Language Neuron selection uses pairwise overlap coefficients of full activation distributions, including negative values, to cluster languages. Experiments on two mLLMs and two held-out corpora show that this method isolates language-specific causal effects more effectively, achieving up to 4.9× higher on-target language damage per neuron while maintaining off-target language performance.
arXiv:2601. 18699v2 Announce Type: replace Abstract: Sequential fine-tuning of Large Language Models (LLMs) adaptation to target tasks often triggers catastrophic forgetting, where the acquisition of novel target skills degrades ancestral capabilities.
arXiv:2608. 05164v1 Announce Type: cross Abstract: Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but whether this geometric similarity has functional consequences for cross-model behavioural control remains untested.
arXiv:2605. 05983v2 Announce Type: replace Abstract: Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones.