arXiv AI By Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

Read the original on arXiv AI →

The paper introduces Fair-ASR, a protocol for evaluating black‑box jailbreak attacks using a shared target‑call budget, addressing the bias of prior studies that rely solely on attack success rate. Re‑evaluating 11 attacks under this protocol shows significant shifts in rankings and highlights that many methods are not efficient in both target and attacker calls. The authors then present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑target‑call budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 18

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

The paper introduces Fair-ASR, a new evaluation protocol for black-box jailbreak attacks that uses shared target-call budgets to provide a fair comparison across methods. Re‑evaluating 11 attacks under this protocol shows that rankings shift significantly with different budgets, and that simple perturbations and templates remain competitive. The authors also present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑call budget.

arXiv Computation and Language
Sep 21

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

The paper presents a systematic study of combining defenses against jailbreak attacks on Large Language Models across different pipeline stages. It introduces a standardized evaluation framework that defines attack-success-rate, controls query budgets, and applies explicit fairness rules. Across 19 attacks and 15 defenses, the study finds that no single defense dominates, but carefully selected combinations can provide strong safety with minimal loss of utility, offering practical guidance for layered defense pipelines.

By Jiale Luo, Eric Han