The paper introduces Fair-ASR, a new evaluation protocol for black-box jailbreak attacks that uses shared target-call budgets to provide a fair comparison across methods. Re‑evaluating 11 attacks under this protocol shows that rankings shift significantly with different budgets, and that simple perturbations and templates remain competitive. The authors also present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑call budget.
The paper presents a systematic study of combining defenses against jailbreak attacks on Large Language Models across different pipeline stages. It introduces a standardized evaluation framework that defines attack-success-rate, controls query budgets, and applies explicit fairness rules. Across 19 attacks and 15 defenses, the study finds that no single defense dominates, but carefully selected combinations can provide strong safety with minimal loss of utility, offering practical guidance for layered defense pipelines.
By Jiale Luo, Eric Han
arXiv:2602. 24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols.
By Zhicheng Fang, Jingjie Zheng, Chenxu Fu, Wei Xu
arXiv:2606. 11409v1 Announce Type: cross Abstract: Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly.
By Malikeh Ehghaghi, Bogl\'arka Ecsedi, Marsha Chechik, Colin Raffel
arXiv:2502. 09755v4 Announce Type: replace-cross Abstract: Safety-aligned LLMs respond to prompts with either compliance or refusal, each corresponding to distinct directions in the model's activation space.
By Amit Levi, Rom Himelstein, Yaniv Nemcovsky, Avi Mendelson, Chaim Baskin
arXiv:2606. 05609v1 Announce Type: cross Abstract: As large language models (LLMs) are widely deployed, identifying their vulnerability through jailbreak attacks becomes increasingly critical.
By Seungwon Jeong, Jiwoo Jeong, Hyeonjin Kim, Yunseok Lee, Woojin Lee