arXiv Computation and Language By Yanxu Mao, Peipei Liu, Tiehan Cui, Zhaoteng Yan, Congying Liu, Datao You

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

Read the original on arXiv Computation and Language →

The paper introduces JMLLM, a multimodal jailbreaking approach that targets text, visual, and auditory inputs to expose vulnerabilities in large language models. It also presents TriJail, a new dataset containing jailbreak prompts across all three modalities. Experiments on TriJail and AdvBench show higher attack success rates and lower time overhead compared to existing methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 23

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

arXiv:2607. 19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates.

By Qingjia Huang, Jingyu Zhang, Jianguo Wu, Yakai Li, Weijuan Zhang, Yankai Rong, Junyi Yao, Shengzhi Zhang, Xiaoqi Jia
arXiv AI
Aug 20

`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs

The paper introduces an adaptive jailbreak attack framework that evaluates both cascaded pipelines and end‑to‑end large audio‑language models (LALMs) under a unified setting. It employs a feedback‑guided mutation engine to automatically generate and refine jailbreak candidates across textual prompts and audio perturbations, thereby broadening attack diversity. Experiments on six audio‑based systems show that both paradigms remain highly vulnerable, with the framework achieving higher attack success rates than existing methods.

By Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
arXiv AI
Jun 11

JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization

arXiv:2606. 11425v1 Announce Type: cross Abstract: Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many target queries.

By Ge Shi, Jun Yin, Donglin Xie, Fangyi Liu, Yucan Li, Menglin Liu