arXiv AI

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking

arXiv:2604. 09222v2 Announce Type: replace-cross Abstract: Audio Large Language Models (ALLMs) enable spoken interaction but introduce new jailbreak vulnerabilities.

arXiv AI
Aug 20

`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs

The paper introduces an adaptive jailbreak attack framework that evaluates both cascaded pipelines and end‑to‑end large audio‑language models (LALMs) under a unified setting. It employs a feedback‑guided mutation engine to automatically generate and refine jailbreak candidates across textual prompts and audio perturbations, thereby broadening attack diversity. Experiments on six audio‑based systems show that both paradigms remain highly vulnerable, with the framework achieving higher attack success rates than existing methods.

By Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
arXiv AI
Jul 23

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

arXiv:2607. 19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates.

By Qingjia Huang, Jingyu Zhang, Jianguo Wu, Yakai Li, Weijuan Zhang, Yankai Rong, Junyi Yao, Shengzhi Zhang, Xiaoqi Jia
arXiv Computation and Language
Sep 21

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

The paper presents a systematic study of combining defenses against jailbreak attacks on Large Language Models across different pipeline stages. It introduces a standardized evaluation framework that defines attack-success-rate, controls query budgets, and applies explicit fairness rules. Across 19 attacks and 15 defenses, the study finds that no single defense dominates, but carefully selected combinations can provide strong safety with minimal loss of utility, offering practical guidance for layered defense pipelines.

By Jiale Luo, Eric Han
arXiv AI
Aug 20

Jailbreaking in the Haystack

The paper "Jailbreaking in the Haystack" introduces NINJA, a jailbreak technique that exploits long-context language models by appending benign, model-generated content to harmful user goals. It demonstrates that the position of harmful goals within the context is crucial for safety, and shows that NINJA significantly boosts attack success rates on models such as LLaMA, Qwen, Mistral, and Gemini. Unlike previous methods, NINJA is low-resource, transferable, less detectable, and compute‑optimal, revealing that carefully crafted benign long contexts can expose fundamental vulnerabilities in modern LMs.

By Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan
arXiv Computation and Language
Sep 17

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

The paper introduces JMLLM, a multimodal jailbreaking approach that targets text, visual, and auditory inputs to expose vulnerabilities in large language models. It also presents TriJail, a new dataset containing jailbreak prompts across all three modalities. Experiments on TriJail and AdvBench show higher attack success rates and lower time overhead compared to existing methods.

By Yanxu Mao, Peipei Liu, Tiehan Cui, Zhaoteng Yan, Congying Liu, Datao You