arXiv Machine Learning By Fengwei Tian, Ravi Tandon

Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks

Read the original on arXiv Machine Learning →

The paper introduces an information‑theoretic framework for studying intent‑hiding jailbreaks against large language models. It defines prior‑posterior matching, where auxiliary tasks are chosen so that the average probability of harmful intent in a bundle matches the overall prior, thereby concealing a harmful target. The authors analyze both query‑independent and query‑dependent settings, proving computational hardness for exact matching, deriving a water‑filling solution for fractional weights, and evaluating the trade‑off between bundle size and target preservation across several open‑source models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 3

Context Inference Attacks Without Jailbreaks

The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.

By Prince Jha, Samuele Poppi, Nils Lukas
arXiv AI
Sep 7

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.

By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv AI
Jun 3

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

arXiv:2606. 02640v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to iteratively refine prompts toward harmful goals.

By Huanli Gong, Zhipeng Wei, Yu Fu, Haz Sameen Shahgir, Ananya Gupta, Yue Dong, N. Benjamin Erichson