arXiv Computation and Language
Sep 3

GAPS: Dimension-Level Gates for Conditional Activation Steering

The paper introduces GAPS, a dimension‑level gating approach for activation steering in language models. GAPS uses two training‑free gates—a static separability gate based on AUROC and a dynamic posterior gate based on a Gaussian model—to selectively apply steering vectors only to neurons that carry reliable concept information or are currently mis‑activated. Experiments on Gemma‑3 and Qwen‑3 show that GAPS improves or matches the performance of token‑level methods, notably reducing Gemma‑3’s toxicity rate from 6.52% to 0.48% under a fixed capability budget.

By Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad, A. B. Siddique