The paper evaluates six automated jailbreak evaluators—HarmBench, JailbreakBench, JailbreakRadar, StrongReject, JADES, and JailMeter—using human-labeled data from JailbreakQR and JailMeter-Eva. It measures each evaluator’s agreement with human judgments, error types, and consistency across attack families, controlling for model-specific variation by using a shared LLM judge where needed. The study finds that JADES performs best overall, with HarmBench and StrongReject also showing strong performance.
By Yujie Mu
arXiv:2608.21895v1 Announce Type: cross
Abstract: Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present in API-serv...
By Aaditya Pratap, Harsh Kasyap, Somanath Tripathy
AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.
By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv:2506. 22666v3 Announce Type: replace-cross Abstract: The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabilities in real-world settings.
By Anamika Lochab, Lu Yan, Patrick Pynadath, Xiangyu Zhang, Ruqi Zhang
arXiv:2609.05850v1 Announce Type: cross
Abstract: Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has reve...
By Quoc Viet Vo, Trung Le, Damith C. Ranasinghe, Ehsan Abbasnejad
The paper introduces SEAV, a verification‑centric framework for evaluating jailbreak attempts against large language models. SEAV decomposes responses into ordered steps and checks both validity and correctness using LLM‑as‑a‑judge and retrieval‑grounded verification. The method reduces false positives by 14.9 percentage points on a strategic‑dishonesty diagnostic and reclassifies 22.1–51.0% of previously successful jailbreaks as invalid across multiple benchmarks.
By Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran