arXiv Machine Learning

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.

arXiv AI
Aug 20

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.

By Godwin Abuh Faruna