GAPS: Dimension-Level Gates for Conditional Activation Steering
Read the original on arXiv Computation and Language →The paper introduces GAPS, a dimension‑level gating approach for activation steering in language models. GAPS uses two training‑free gates—a static separability gate based on AUROC and a dynamic posterior gate based on a Gaussian model—to selectively apply steering vectors only to neurons that carry reliable concept information or are currently mis‑activated. Experiments on Gemma‑3 and Qwen‑3 show that GAPS improves or matches the performance of token‑level methods, notably reducing Gemma‑3’s toxicity rate from 6.52% to 0.48% under a fixed capability budget.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.