The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.
By Prince Jha, Samuele Poppi, Nils Lukas
arXiv:2609.06540v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role...
By Qi Wang, Chengcheng Wan, Jiangtao Wang
AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.
By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv:2605.27110v2 Announce Type: replace-cross
Abstract: In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that elicits malicious information through in...
By Xuan Luo, Yue Wang, Geng Tu, Jing Li, Ruifeng Xu
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
arXiv:2606. 02640v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to iteratively refine prompts toward harmful goals.
By Huanli Gong, Zhipeng Wei, Yu Fu, Haz Sameen Shahgir, Ananya Gupta, Yue Dong, N. Benjamin Erichson