arXiv:2608.29109v1 Announce Type: new
Abstract: Large language models often answer structurally unanswerable questions, such as computing cot(-540{\deg}) or evaluating (1).startswith("1"), instead of...
By Yucheng Du, Xiyang Hu
arXiv:2607. 08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations.
By Ege \c{C}akar, Hannah Guan, Kayden Kehe
arXiv:2605. 21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations.
By Giorgio Piras, Raffaele Mura, Fabio Brau, Maura Pintor, Luca Oneto, Fabio Roli, Battista Biggio
arXiv:2603.27518v4 Announce Type: replace
Abstract: Aligned language models that are trained to refuse harmful requests also exhibit over-refusal: they decline safe instructions that seemingly resemb...
By Utsav Maskey, Mark Dras, Usman Naseem
arXiv:2608. 11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters.
By Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari
arXiv:2606. 28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood.
By Yanchen Yin, Dongqi Han, Linghui Li
arXiv:2607. 20473v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests.
By Yeonjea Kim, Bumjin Park, Jaesik Choi
arXiv:2603. 21396v5 Announce Type: replace Abstract: Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness.
By Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey
arXiv:2606. 22686v2 Announce Type: replace-cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque.
By Shivam Ratnakar, Kartikeya Vats
arXiv:2606. 04413v1 Announce Type: new Abstract: Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would be an obstacle.
By Mohammad Omar Khursheed, Baram Sosis, Fabien Roger
arXiv:2608. 08027v1 Announce Type: cross Abstract: Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data.
By Laiqiao Qin, Tianqing Zhu, Longxiang Gao, Wanlei Zhou
arXiv:2605.01913v2 Announce Type: replace-cross
Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vul...
By Sadia Asif, Mohammad Mohammadi Amiri