arXiv:2608. 12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks.
By Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei
The study investigates how category‑specific harmfulness signals, isolated by removing the shared general harmfulness component, influence large language model (LLM) safety. Using activation steering across 11 risk categories in three instruction‑tuned LLMs, the authors find that the presence of harmfulness in these category residuals varies by category and that the pattern of inducing refusal is even more model‑dependent. Additionally, category residuals were shown to enhance the models’ downstream alignment with the shared general harmfulness representation, indicating that fine‑grained signals play a role beyond the general component.
By Soyeon Park (KAIST), Seogyeong Jeong (KAIST), Sunwoo Kim (KAIST), Alice Oh (KAIST)
arXiv:2606. 05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content.
By Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu, Wenxuan Zhang
arXiv:2606. 02423v1 Announce Type: cross Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve harmful outcomes beyond their capabilities through extended interactions.
By Ruohao Guo, Wei Xu, Alan Ritter
arXiv:2608. 08383v1 Announce Type: cross Abstract: Steering vectors are a lightweight tool for controlling LLM behavior.
By Yuxiao Li, Gjergji Kasneci
arXiv:2607. 02079v1 Announce Type: cross Abstract: We present HaloGuard 1.
By Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin