arXiv AI

Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs

arXiv:2602. 00945v2 Announce Type: replace-cross Abstract: LLMs are multilingual by training, yet their lingua franca is often English, reflecting English language dominance in pretraining.

arXiv Machine Learning
Aug 19

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.

By Sihan Wang, Jiayi Zhao, Qingyan Cao, Hongbo Yao, Lin Shu
arXiv Computation and Language
Sep 11

Distribution-aware Language Neuron Identification in Multilingual Large Language Models

The paper introduces a new method for identifying language-specific neurons in multilingual large language models (mLLMs). Unlike previous entropy-based approaches that only consider positive activations, the proposed Distribution-aware Language Neuron selection uses pairwise overlap coefficients of full activation distributions, including negative values, to cluster languages. Experiments on two mLLMs and two held-out corpora show that this method isolates language-specific causal effects more effectively, achieving up to 4.9× higher on-target language damage per neuron while maintaining off-target language performance.

By Minjun Kim, Inho Won, Junghun Yuk, Dongyeon Kim, Jihyo Kim, KyungTae Lim
arXiv Machine Learning
Jun 15

Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions

arXiv:2605. 05983v2 Announce Type: replace Abstract: Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones.

By Yuntai Bao, Qinfeng Li, Xinyan Yu, Ge Su, Wenqi Zhang, Liu Yan, Haiqin Weng, Jianwei Yin, Xuhong Zhang