arXiv Machine Learning By Nima Eshraghi, Lovedeep Gondara, Yuqing Huang, Sagarika Suresh, Leizer Teran, Jithin Pradeep, Xiaotong Xu, Fanny Chevalier

A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2607. 05615v1 Announce Type: new Abstract: Activation steering via sparse autoencoders (SAEs) enables behavioral control of large language models without task-specific fine-tuning, but standard methods apply the steering signal at every generated token, incurring constant per-token perturbation that risks degrading fluency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Minimally Invasive Steering of Language Models

The paper introduces Minimally Invasive Steering Vector Optimization (MISVO), a method that adjusts a frozen language model’s final hidden states by adding vectors to steer outputs toward a test‑time reward while penalizing changes using the local KL geometry of the token distribution. MISVO derives an analytic Fisher term and a suffix score‑function term, showing that the suffix term is second‑order and that Fisher surrogates match the full KL gradient to first order. Experiments on preference and code‑generation tasks with 1B–14B parameter models demonstrate that MISVO achieves the highest mean reward in most settings while maintaining diversity and coherence comparable to Best‑of‑N.

By Taha Entesari, Jingyu Zhang, Daniel Khashabi, Mahyar Fazlyab
arXiv Machine Learning
Jun 15

Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions

arXiv:2605. 05983v2 Announce Type: replace Abstract: Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones.

By Yuntai Bao, Qinfeng Li, Xinyan Yu, Ge Su, Wenqi Zhang, Liu Yan, Haiqin Weng, Jianwei Yin, Xuhong Zhang
arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
arXiv Computation and Language
Sep 16

There Is More to Refusal in Large Language Models than a Single Direction

The paper challenges the notion that refusal in large language models is governed by a single direction in activation space. It demonstrates that different refusal and non‑compliance categories map to distinct geometric directions, yet steering along any of these directions yields similar refusal–over‑refusal trade‑offs, acting as a shared one‑dimensional control knob. Using sparse autoencoders, the authors reveal a structured internal representation of refusal, comprising a reusable core of shared latents and style‑ or domain‑specific latents, and show that linear interventions collapse this structure into uniform behavioral control.

By Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, Husrev Taha Sencar