arXiv:2607. 10226v1 Announce Type: new Abstract: We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior.
By Daming Luo
arXiv:2606. 08365v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across contexts and perturb unrelated features.
By Evan Duan
arXiv:2604. 26866v2 Announce Type: replace-cross Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction.
By Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas
arXiv:2607. 19364v2 Announce Type: replace Abstract: Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning.
By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
arXiv:2608. 03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs.
By Yu Feng, Chunting Zang, Chen Shen, Rui Miao, Ge Teng, Weidong Cai, Jieping Ye
arXiv:2604. 23130v2 Announce Type: replace-cross Abstract: Jailbreak attacks expose a persistent failure mode in safety-aligned LLMs: models can be pushed into harmful behavior, but the internal representations enabling this shift remain poorly localized.
By Nilanjana Das, Mathew Dawit, Aman Chadha, Manas Gaur