arXiv Computation and Language

There Is More to Refusal in Large Language Models than a Single Direction

The paper challenges the notion that refusal in large language models is governed by a single direction in activation space. It demonstrates that different refusal and non‑compliance categories map to distinct geometric directions, yet steering along any of these directions yields similar refusal–over‑refusal trade‑offs, acting as a shared one‑dimensional control knob. Using sparse autoencoders, the authors reveal a structured internal representation of refusal, comprising a reusable core of shared latents and style‑ or domain‑specific latents, and show that linear interventions collapse this structure into uniform behavioral control.

arXiv Machine Learning
Jul 8

A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

arXiv:2607. 05615v1 Announce Type: new Abstract: Activation steering via sparse autoencoders (SAEs) enables behavioral control of large language models without task-specific fine-tuning, but standard methods apply the steering signal at every generated token, incurring constant per-token perturbation that risks degrading fluency.

By Nima Eshraghi, Lovedeep Gondara, Yuqing Huang, Sagarika Suresh, Leizer Teran, Jithin Pradeep, Xiaotong Xu, Fanny Chevalier
arXiv Machine Learning
Aug 27

Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

The paper investigates how refusal training shapes the internal geometry of language models, showing that activation updates from refusal-completion losses create a distinct low‑dimensional refusal subspace. In a case study on OLMo‑2‑0425‑1B‑Instruct, the authors link the brittleness of refusal directions to repetitive refusal prefixes and demonstrate that using diverse refusal starts can increase the stable rank of gradients, thereby hardening the model against vector‑ablation attacks. The work provides insights into the emergence of safety‑critical features and offers a potential strategy to strengthen refusal robustness.

By Andrey Labunets
Hugging Face Trending Papers
Aug 13

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.

arXiv Machine Learning
Sep 10

Disentangling Steering Vectors

arXiv:2609.07037v1 Announce Type: new Abstract: Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional...

By Takeru Hiramatsu, Kyohei Atarashi, Koh Takeuchi, Hisashi Kashima