arXiv AI

Divergent Response Modes in Frontier Language Models Under Steering Pressure

arXiv:2608. 06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines.

arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou