Divergent Response Modes in Frontier Language Models Under Steering Pressure
arXiv:2608. 06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines.
arXiv:2509. 13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets.
arXiv:2608. 06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines.
arXiv:2609.06289v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...
arXiv:2610.00601v1 Announce Type: cross Abstract: Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface fo...
arXiv:2608. 08383v1 Announce Type: cross Abstract: Steering vectors are a lightweight tool for controlling LLM behavior.
arXiv:2606. 22686v2 Announce Type: replace-cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque.
The paper introduces CAM-Steer, a Category‑Adaptive Multi‑category Safety Steering framework that estimates risk for each harm category by comparing hidden states to safe and unsafe prototypes. It then combines safety directions into a single steering vector and applies a rotation whose angle is set by the estimated risks, preserving the hidden‑state norm. Experiments on three LLM backbones and seven harm categories show that CAM‑Steer outperforms baselines in defense success rate, even when multiple harm categories co‑occur, with negligible inference overhead.
arXiv:2608. 11227v1 Announce Type: new Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining.
arXiv:2510. 04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactions in socially interactive settings.
arXiv:2609.06951v1 Announce Type: cross Abstract: Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activa...
arXiv:2606. 05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content.
arXiv:2608. 13250v1 Announce Type: cross Abstract: Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge.
arXiv:2607. 11871v1 Announce Type: cross Abstract: Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations.