arXiv Machine Learning

Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks

The paper introduces an information‑theoretic framework for studying intent‑hiding jailbreaks against large language models. It defines prior‑posterior matching, where auxiliary tasks are chosen so that the average probability of harmful intent in a bundle matches the overall prior, thereby concealing a harmful target. The authors analyze both query‑independent and query‑dependent settings, proving computational hardness for exact matching, deriving a water‑filling solution for fractional weights, and evaluating the trade‑off between bundle size and target preservation across several open‑source models.

arXiv Machine Learning
Sep 3

Context Inference Attacks Without Jailbreaks

The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.

By Prince Jha, Samuele Poppi, Nils Lukas
arXiv AI
Sep 7

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.

By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv AI
Jun 3

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

arXiv:2606. 02640v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to iteratively refine prompts toward harmful goals.

By Huanli Gong, Zhipeng Wei, Yu Fu, Haz Sameen Shahgir, Ananya Gupta, Yue Dong, N. Benjamin Erichson
arXiv Machine Learning
Jun 26

Jailbreaking for the Average Jane: Choosing Optimal Jailbreaks via Bandit Algorithms for Automatically Enhanced Queries

arXiv:2606. 26936v1 Announce Type: cross Abstract: With a profusion of jailbreaks for LLMs now widely known, a growing concern is that non-expert malicious actors ("the average Jane") could elicit actionable responses to malicious requests.

By Prarabdh Shukla, Ritik, Suhas Rao, Arpit Agarwal, Arjun Bhagoji
arXiv AI
Jul 9

NonTextual Target Attack

arXiv:2510. 02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses.

By Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Hongsheng Hu, Xiaojun Jia, Di Wang, Zhan Qin, Kui Ren