arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.
By Sahil Kadadekar
arXiv:2606. 25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy.
By Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee
arXiv:2607. 00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies.
By Shei Pern Chua, Fangzhao Wu
arXiv:2607.15218v2 Announce Type: replace
Abstract: Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become uns...
By Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu
arXiv:2606. 05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content.
By Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu, Wenxuan Zhang