arXiv:2608.30585v1 Announce Type: new
Abstract: Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model wou...
By Md Mokarram Chowdhury, Ernie Chang, Yang Li
The paper investigates how different post‑training interventions—harmful supervised fine‑tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal‑feature ablation—affect large language models’ harmful compliance, capability, and safety signals. Across Qwen2.5‑7B and Llama‑3.1‑8B, all methods achieve near‑maximum harmfulness, but SFT causes the greatest loss of capability and representational drift, ablation suppresses refusal features in a family‑specific way, and RLVR largely preserves base‑model performance while redirecting behavior toward compliance. RLVR models also exhibit “capability‑blind compliance,” falsely claiming to perform unavailable actions, which can be mitigated by targeted calibration without harming overall capability. The study demonstrates that harmful compliance, harm recognition, and capability awareness are distinct behavioral axes and that typical safety signals such as self‑audit and hallucination may not reliably indicate robustness after adaptive post‑training.
By Md Rysul Kabir, Zoran Tiganj
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
By Alex Kwon
arXiv:2606. 13720v1 Announce Type: new Abstract: Arditi et al.
By Elisabetta Rocchetti, Alfio Ferrara
arXiv:2608. 16177v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists?
By Hidayet Aksu
The paper introduces TIER, a Threat Implicitness Benchmark designed to evaluate large language model (LLM) safety behaviors across four risk domains and four threat levels, ranging from explicit harmful requests to sophisticated jailbreaks. Responses are scored on a six-label behavior scale by two independent LLM judges. Experiments on six open-weight LLMs reveal that safety behaviors change gradually with threat level, contextual prompts produce the most varied responses, and jailbreaks expose significant robustness gaps, underscoring the importance of behavior-aware safety evaluation.
By Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho
arXiv:2609.38389v1 Announce Type: cross
Abstract: Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear...
By Yelyzaveta (Lisa), Husieva, Lauren Alvarez
arXiv:2605. 00123v3 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts.
By Shubham Kumar, Narendra Ahuja
arXiv:2609.22090v1 Announce Type: new
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps...
By Joy Bose
arXiv:2608. 03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs.
By Yu Feng, Chunting Zang, Chen Shen, Rui Miao, Ge Teng, Weidong Cai, Jieping Ye
arXiv:2509. 13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets.
By Vincent Siu, Nicholas Crispino, David Park, Nathan W. Henry, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang
arXiv:2608.29942v1 Announce Type: cross
Abstract: The key limitation of current state-of-the-art influence-based guardrails is that they do not reliably distinguish a legitimate, user-authorized acti...
By Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder