arXiv Machine Learning By Tong Zhang, Zexin Li, Simin Chen, Yun Peng

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Read the original on arXiv Machine Learning →

arXiv:2607. 24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 21

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

The paper presents a systematic study of combining defenses against jailbreak attacks on Large Language Models across different pipeline stages. It introduces a standardized evaluation framework that defines attack-success-rate, controls query budgets, and applies explicit fairness rules. Across 19 attacks and 15 defenses, the study finds that no single defense dominates, but carefully selected combinations can provide strong safety with minimal loss of utility, offering practical guidance for layered defense pipelines.

By Jiale Luo, Eric Han
arXiv AI
Sep 7

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.

By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv AI
Aug 5

AI Security Leaderboard: Methodology, Results and Minimal Standard

arXiv:2608. 03070v1 Announce Type: cross Abstract: Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much protection these safeguards provide, or how consistently across developers.

By Jasper Timm, Lukas Struppek, Ziwei Xu, Grace Cheong, Oscar Mata, Dan Zhao, Mick Yang, Isadora De Andrade, Xiaojun Jia, Yiming Li, Samuel Bauer, Heather McIntyre, Adam Gleave, Edward Yee, Kellin Pelrine