arXiv:2601. 22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels.
By Charles Westphal, Keivan Navaie, Fernando E. Rosas
arXiv:2606. 09411v1 Announce Type: cross Abstract: Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs.
By Charles Westphal, Timothy Douglas, Keivan Navaie, Tiago Pimentel, Fernando E. Rosas
Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.
By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov
arXiv:2606. 09135v1 Announce Type: cross Abstract: We demonstrate that widely deployed Large Language Model (LLM) inference stacks harbor a steganographic channel that requires no modification to model weights, sampling code, or output distributions.
By Felix M\"achtle, Jonas Sander, Sebastian Berndt, Ben Weimar, Nils Loose, Thomas Eisenbarth
The paper critiques current encoded‑prompt safety benchmarks that focus only on harmful requests, showing that such tests can misrepresent a model’s safety. By evaluating the benign arm under the same encoding, the authors reveal a substantial drop in the harm gap—sometimes to zero—indicating that the encoding masks true safety deficiencies. Across multiple large models and training pipelines, they document that the encoding can either hide or falsely inflate safety metrics, and they identify twelve specific instrument defects that contribute to these misleading results.
By Haoyu Zhang, Haowen Xu, Xiao Luo, Hanwen Liu, Yang Chen, Zijian Xiao, Yi Feng, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.
By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r
arXiv:2608. 08027v1 Announce Type: cross Abstract: Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data.
By Laiqiao Qin, Tianqing Zhu, Longxiang Gao, Wanlei Zhou
The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.
By Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv:2604. 06247v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or images.
By Guy Azov, Ofer Rivlin, Guy Shtar
arXiv:2608. 01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR).
By Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2606. 18530v1 Announce Type: cross Abstract: Domain-camouflaged injection attacks embed malicious instructions in retrieved content using domain-appropriate vocabulary, evading standard detectors that rely on syntactic injection markers.
By Aaditya Pai
arXiv:2609.36570v1 Announce Type: cross
Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
By Mark Russinovich