arXiv Computation and Language By Qian Chen, Shiliang Xiao, Yuzhi Liang

OASIS: Optimizing Attacker Sequences for Hard-Label Black-Box Text Attacks

Read the original on arXiv Computation and Language →

OASIS is a method for optimizing attacker sequences in hard‑label black‑box text attacks. It first performs a one‑time bi‑objective search over candidate sequences to balance attack success rate and perturbation, then reuses the selected fixed global chain during execution. Experiments on multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 9

NonTextual Target Attack

arXiv:2510. 02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses.

By Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Hongsheng Hu, Xiaojun Jia, Di Wang, Zhan Qin, Kui Ren
arXiv Computation and Language
Sep 21

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

The paper presents a systematic study of combining defenses against jailbreak attacks on Large Language Models across different pipeline stages. It introduces a standardized evaluation framework that defines attack-success-rate, controls query budgets, and applies explicit fairness rules. Across 19 attacks and 15 defenses, the study finds that no single defense dominates, but carefully selected combinations can provide strong safety with minimal loss of utility, offering practical guidance for layered defense pipelines.

By Jiale Luo, Eric Han
arXiv AI
Jun 4

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

arXiv:2606. 04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs.

By Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue, Yue Zhao, Chaowei Xiao
arXiv Machine Learning
Sep 1

Learning diverse attacks on large language models for robust red-teaming and safety tuning

arXiv:2405.18540v3 Announce Type: replace-cross Abstract: Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of larg...

By Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain